Daily Twitter Digest

Computer Science

LLM-Simulated Conversations Underplay Human Miscommunication and Conflict

The paper argues that evaluating LLM-simulated human conversations should account for the fact that real human dialogue is full of misunderstandings, interruptions, and uncooperative moments—not just coherent, cooperative exchanges. The authors introduce CoCoEval, a framework that detects 10 types of inconsistent or uncollaborative behaviors turn-by-turn in simulated professional-scenario conversations. Testing GPT-4.1, GPT-5.1, and Claude Opus 4 against human transcripts, they find LLM-simulated conversations show far fewer such behaviors than real ones under standard prompting, and that prompt engineering or fine-tuning fails to reliably fix this—often overcorrecting into specific artificial behaviors instead. The authors say these gaps are invisible to conventional conversation-level Likert-scale evaluations, raising concerns about using LLMs as stand-ins for human social interaction in research.

Discussion: 2 tweets from 2 authors · @RyoKamoi, @RyoKamoi_ja

Engineering

Position Paper Argues Generalist Robots Need a New Safety Framework

The paper argues that as robots gain generalist capabilities—performing everything from cooking to infrastructure repair—traditional safety notions centered on collision and force avoidance become insufficient. The authors contend that safety must now account for context (e.g., when it's safe to shut off building power), unspoken user intent (e.g., not mixing dangerous cleaning chemicals when asked to "clean the kitchen"), and hard-to-model physical consequences like burning food. They propose "embodied AI safety" as a new paradigm, presenting a taxonomy of emerging risks and a full-stack research agenda spanning the robot's entire lifecycle, arguing that digital (bits) and physical (atoms) safety can no longer be treated separately.

Discussion: 2 tweets from 2 authors · @andrea_bajcsy, @Majumdar_Ani

Biology

New Therizinosaur Species Kayrasaurus Discovered in China's Jehol Biota

This paper (abstract unavailable) reportedly describes a new therizinosaur, Kayrasaurus changliensis, from the Early Cretaceous Jehol Biota of China, and examines ecomorphological patterns among early therizinosaurians based on the paper's title. Therizinosaurs are an unusual group of theropod dinosaurs known for their herbivorous adaptations, long claws, and bizarre body plans, and new Jehol specimens often help clarify their early evolutionary diversification. Twitter discussion centers on a paleoartist's commissioned reconstruction artwork, made in collaboration with the Institute of Vertebrate Paleontology and Paleoanthropology (IVPP), depicting Kayrasaurus alongside three other therizinosaur species from the Jehol Biota. Commentary is primarily celebratory of the artwork and new species rather than critical, with no substantive skepticism raised in the visible discussion.

Discussion: 1 tweets from 1 authors · @_HETAKA

Medicine

Powered vs. Manual Toothbrushes Show Similar Results for People With Intellectual Disability

This systematic review and meta-analysis examined whether powered or manual toothbrushes work better for supervised toothbrushing among people with mild-to-moderate intellectual disabilities. Drawing on 11 studies from seven countries (with only three suitable for meta-analysis of plaque reduction via the Quigley-Hein index), the authors found powered and manual toothbrushes to be comparably effective overall, though five of the eleven studies reported powered brushes as superior. The authors emphasize that caregiver supervision itself appears to be the key driver of improved oral hygiene, regardless of toothbrush type, while cautioning that the small pooled sample (n=3 studies) limits certainty.

Discussion: 1 tweets from 1 authors · @afcp_01

Computer Science

Cooperation Emerges When Computation and Reproduction Share One Energy Budget

The paper introduces 'Autopoietic Game Theory,' a computational model where social interactions, self-replication, and computation costs are all endogenous and co-evolve, rather than being studied in isolation as in classical evolutionary game theory or artificial life. Using randomly initialized Z80 machine-code programs as a substrate, the authors show that embedding a social dilemma directly into the physics of computation—where defection ('stealing') destroys shared energy and slows execution—causes cooperative, self-replicating strategies to dominate, especially under resource scarcity. They report that evolved programs suppress stealing across several test environments, with spatial structure further boosting complexity and task performance, and that the framework extends to exogenous tasks like math problems framed as sequential social dilemmas. The author's Twitter thread frames the core finding as evidence that self-interested, self-improving, self-replicating agents can learn to cooperate 'from scratch' once social behavior, computation, and reproduction draw from a single shared energy budget, positioning this as a novel bridge between evolutionary game theory and artificial life research.

Discussion: 1 tweets from 1 authors · @kjha02

Computer Science

Coverage, Not Cross-Entropy, Explains Why Pre-Training Enables Post-Training

The paper argues that cross-entropy loss, the standard metric for pre-training quality, is a poor proxy for how well a model will perform after fine-tuning or test-time scaling (e.g., Best-of-N). Instead, the authors propose 'coverage' — the probability mass a pre-trained model assigns to high-quality responses — as the key quantity that is both necessary and sufficient for post-training success. They show theoretically that next-token prediction implicitly drives models toward good coverage, that coverage generalizes faster than cross-entropy (avoiding spurious dependence on sequence length), and they derive practical interventions (checkpoint selection, gradient normalization, decoding strategies) with provable coverage benefits. On Twitter, the authors highlighted this as a follow-up to earlier work questioning whether cross-entropy SFT is the right way to prepare models for RL, framing the new 'TailSFT' method as a lightweight, theoretically grounded way to boost coverage and improve post-RL performance, since RL compute is expensive and every pre-training/SFT step should count.

Discussion: 1 tweets from 1 authors · @SadhikaMalladi

Computer Science

OpenAI's VPT Learns Minecraft by Watching Unlabeled YouTube Videos

The paper introduces Video PreTraining (VPT), a method for training sequential decision-making agents on internet-scale unlabeled video by first learning an inverse dynamics model from a small labeled dataset, then using it to label vast troves of online Minecraft gameplay footage. This labeled data trains a behavioral prior using native human controls (mouse and keyboard at 20Hz) that shows nontrivial zero-shot skill and can be fine-tuned via imitation learning and RL to solve hard-exploration tasks—including crafting diamond tools, a feat that normally takes skilled human players over 20 minutes of gameplay. The authors claim this is the first time an agent has achieved this benchmark, and that some tasks reach human-level performance.

Discussion: 1 tweets from 1 authors · @jetsetworm

Physics

Mutual Information Dynamics Recovers KS Entropy of Chemical Chaos

The paper studies open chemical reaction networks whose finite-size dynamics are stochastic but whose macroscopic (thermodynamic) limit obeys deterministic rate equations that can be chaotic. The authors show theoretically that a time-dependent rate of information loss, built from two-time mutual information in the stochastic description, converges to the Kolmogorov-Sinai entropy of the deterministic chaotic dynamics as the macroscopic limit is taken, and they verify this numerically for a three-species, seven-reaction Markov jump process. This is presented as the final installment of a three-part 'mutual information dynamics' series.

Discussion: 1 tweets from 1 authors · @sasa3341

Biology

Rare White Dryinid Wasp Identified as Parasitoid of Invasive Okra Pest in Japan

Researchers surveying okra and Malvaviscus plants in southern Kyushu in 2025 report the first confirmed natural-enemy record linking the dryinid wasp Aphelopus nivealis to nymphs and adults of Amrasca biguttula, a leafhopper pest spreading through Asia, Africa, and the Americas. Using morphology plus COI DNA barcoding, the authors matched local specimens to Chinese A. nivealis populations (0.0–1.1% divergence) and also described a new, genetically distinct species, Aphelopus obesulus, from a single male, expanding the known diversity of this parasitoid genus and its potential use in biocontrol. A tweet in Japanese highlighted the finding as notable because A. nivealis is known as an unusually pure-white dryinid, making its identification as a natural enemy of the increasingly damaging okra pest particularly striking to entomology followers.

Discussion: 1 tweets from 1 authors · @buruninja

Computer Science

Do LLMs Internalize Human Values or Just Mimic Outputs? A Steering-Vector Test

Based on Twitter discussion only (no abstract available), this paper reportedly tests whether large language models merely produce value-consistent outputs or actually encode relationships between values internally, using Schwartz's theory of basic human values (20 categories arranged in a circular structure of compatible/conflicting values). From ~26,000 filtered question-value-answer quadruples, the authors build per-value "steering vectors" and check whether the geometric relationships among 20 values in the model's internal representations match the theoretical circular structure, rather than just measuring single-answer accuracy. The tweet highlights a contrast between behavior-centric methods (like OPT, which directly optimizes for desired outputs) and distribution-based methods (like SAS, which derives directions from sparse features in activation distributions across affirmative/negative answers), reporting a much higher "Theory Rank Correlation" for SAS (0.51 on Qwen3.5-9B-Base) versus OPT (0.11), suggesting SAS better preserves the theoretical value structure internally.

Discussion: 1 tweets from 1 authors · @itarutomy

Mathematics

AI Research Agent Claims Proof of Komlós and Beck-Fiala Conjectures

The paper claims a $3\sqrt{2\pi}$ bound for the Komlós signing problem, showing any finite family of unit-norm-bounded real vectors admits a signed sum with bounded $\ell_\infty$-norm independent of dimension or family size. It also derives a consequence for the Beck-Fiala conjecture, giving a two-coloring bound with the conjectured square-root dependence on set membership degree $t$, via a technical result on directional total variation and a 'Banaszczyk transform' for convex sets. Notably, the authors state the proof itself was discovered by an automated AI research agent called Odin.

Discussion: 1 tweets from 1 authors · @AlgoSvensson

Biology

New Volume Pushes Biosemiotics as Framework for Theoretical Biology

This edited MIT Press volume gathers 25 biologists and philosophers—including Denis Noble, Scott Gilbert, and Stuart Kauffman—to advance biosemiotics, the study of sign processes and meaning-making within living systems, as an explanatory framework for theoretical biology. Positioned as a successor to Waddington's influential 1968-72 'Towards a Theoretical Biology' series, it aims to identify the most pressing conceptual problems in biology and propose how sign-based approaches might address them. The Twitter discussion simply flagged the book's open-access availability, noting its lineage from Waddington's field-shaping collection and its roster of prominent contributors, without offering substantive critique.

Discussion: 1 tweets from 1 authors · @adammcroom

Biology

Preprint Identifies Island-Specific Feline Leukemia Virus Lineage on Kikaijima

A preprint (no abstract available, so details are based on the tweet discussion only) reports the identification of a genetically distinct, island-specific lineage of Feline Leukemia Virus (FeLV) circulating among cats on Kikaijima Island in Kagoshima Prefecture, Japan. The work appears to stem from viral genetic surveillance of local cat populations, suggesting geographic isolation may have allowed a unique FeLV strain to evolve or persist there. The Twitter discussion is limited to an announcement from the author's lab celebrating the preprint's release, authored by a first-year PhD student, with no substantive scientific critique offered yet.

Discussion: 1 tweets from 1 authors · @saiaka3

Social Science

Book Argues Reason Is 'Tethered' to Older Brain Systems Driving Feelings and Behavior

In 'Reason and Less,' Vinod Goel proposes that human reasoning doesn't float freely above biology but is tethered to evolutionarily older autonomic, instinctive, and associative brain systems, with feelings generated in ancient brainstem structures serving as the common currency that lets these systems interact to produce behavior. Drawing on neuroscience, psychology, philosophy, and evolutionary biology, Goel argues this 'tethered rationality' model implies that behaviors like climate denial, racism, or political tribalism can't be changed by simply targeting beliefs, since they're rooted in this deeper feeling-based control structure that maximizes pleasure and minimizes displeasure. The book is being highlighted as freely available open access.

Discussion: 1 tweets from 1 authors · @adammcroom

Social Science

Study Tracks Where Political Science's Most-Cited Papers Get Published

Using Web of Science citation data, this study identifies the most-cited political science journal articles by decade from the 1940s to the 2020s. It finds that while the American Political Science Review remains the leading outlet, its dominance has waned as highly cited work now appears across a broader range of journals, and methodological papers have come to account for a growing share of the discipline's most influential research. Twitter commentary highlighted the visualization mapping this shift, noting APSR's continued but diminished lead, the diversification of influential outlets, and the persistent dominance of US-based journals in the field.

Discussion: 1 tweets from 1 authors · @Wubenmensheng

Biology

Flightless Stick Insects Show Signs of Bird-Mediated Long-Distance Dispersal

Researchers used mitochondrial COI haplotypes, nuclear SSR markers, and genome-wide SNPs to study population structure of the flightless stick insect Ramulus mikado across Japan. They found low, mostly non-significant genetic differentiation and identical or near-identical genotypes in populations separated by large distances—patterns hard to explain by the insect's limited walking range alone. The authors argue this supports earlier experimental findings that stick insect eggs can survive passage through bird digestive tracts, making avian-mediated passive dispersal the most plausible explanation for the observed phylogeographic pattern. The Twitter discussion was minimal, mainly a share noting the paper is open access, without substantive critique offered in the visible commentary.

Discussion: 1 tweets from 1 authors · @tugutuguk

Medicine

Study Examines How Pulsed Field Ablation Affects Vein of Marshall Signals

No abstract was available for this paper, so this summary is based on discussion only. The study reportedly examines whether pulsed field ablation (PFA) used for pulmonary vein isolation also modifies electrograms in the Vein of Marshall (VoM), a structure implicated in atrial fibrillation recurrence. One tweet summarizing the findings (from Nakatani et al., JACC Clin Electrophysiol 2026) suggests that while VoM modification during PFA appears associated with better outcomes, achieving it intentionally via PFA is difficult, implying that targeted ethanol infusion into the VoM remains an important complementary treatment option rather than being replaced by PFA alone.

Discussion: 1 tweets from 1 authors · @yosuke3gbst

Mathematics

Paper Argues Statistics Needs Proof Assistants Like Lean

The paper presents no new theoretical results; instead it offers five case studies showing how classical statistics problems can be formalized and checked with proof-assistant software. The authors argue that phrases like "under the usual regularity conditions" hide substantial mathematical content that machine formalization forces to be made explicit, and they call for community-built repositories of formalized statistical axioms and theorems, plus a rethink of graduate training around these tools, drawing an analogy to how high-level programming languages transformed empirical research. Twitter discussion (a single tweet with modest engagement) simply flagged the paper as relevant to the Lean/interactive-theorem-proving community, with no substantive critique surfacing.

Discussion: 1 tweets from 1 authors · @Jose_A_Alonso

Physics

Zenodo Preprint Claims Algebraic Origin for an '86↔86' Dirac-Fock Symmetry

No abstract is available for this Zenodo-hosted paper, so the summary below is based only on its title and the author's tweets. The work reportedly derives, via index-theoretic and algebraic arguments, a claimed '86 ↔ 86 capacity symmetry' in the relativistic Dirac-Fock vacuum, presenting corrected field arithmetic and what the authors call 'numerology guards' to distinguish genuine structure from coincidental numerical patterns. Twitter discussion consists mainly of the author's own announcements of the paper and an updated version, with no independent commentary, criticism, or verification visible in the available tweets. Given the self-promotional nature of the posts and the paper's own inclusion of 'numerology guards' in its title, some readers may want to treat the claimed symmetry with caution pending peer review.

Discussion: 2 tweets from 1 authors · @JCPEREZCODEX, @JCPEREZCODEX

Medicine

Meta-Analysis Finds Neutropenic Diet Offers No Broad Infection Benefit

This systematic review and meta-analysis pooled 15 studies (4,320 patients) comparing a neutropenic diet (ND) to a liberalized diet (LD) in cancer patients, finding no significant differences in major infections, pneumonia, diarrhea, or mortality overall. However, in RCTs restricted to hematologic malignancy/HSCT patients (6 studies, 855 patients), ND was associated with lower bacteremia/fungemia risk (RR 0.67), while other outcomes remained unchanged. The authors conclude that routine ND is not supported for most cancer patients, though selected high-risk HSCT or prolonged-neutropenia patients may still warrant cautious dietary practices.

Discussion: 1 tweets from 1 authors · @DraMartinezLago