Daily Twitter Digest

Chemistry

Chiral Dopant Induces Spin-Selective Charge Transport in Polymer LEDs

Researchers doped an achiral polymer (PFO) with a chiral perylene diimide, using thermal annealing to induce supramolecular chirality that produces circular dichroism and circularly polarized luminescence. Devices made from these films show circularly polarized electroluminescence with dissymmetry factors around 10⁻³, while magnetic conductive AFM measurements reveal spin-selective carrier transport consistent with the chiral-induced spin selectivity (CISS) effect, with opposite spin polarization for each enantiomer. The authors argue this points to a simple strategy for adding spin-selective transport and circularly polarized light emission to otherwise achiral polymer semiconductors, potentially involving spin-polarized carrier recombination alongside conventional chirality-induced emission.

Discussion: 2 tweets from 2 authors · @kitasato_kinou, @hasegawa_masasi

Computer Science

Beyond Accuracy: What to Measure When a Benchmark Saturates

The paper argues that retiring saturated benchmarks in favor of harder ones misses valuable evaluation opportunities, and uses CORE-Bench Hard (a computational reproducibility benchmark for AI agents) as a case study. The authors surface construct-validity issues (e.g., shortcuts) invisible with weaker agents, introduce updated CORE-Bench v1.1 and an out-of-distribution suite, and show that even after accuracy saturates, the benchmark remains useful for measuring efficiency, reliability, and model-versus-scaffold contributions. A small randomized experiment also finds human-agent collaboration roughly doubles speed on real reproducibility tasks, likely an underestimate since some human-only trials hit the time limit. Together they propose a more multidimensional alternative to accuracy-centric agent evaluation.

Discussion: 1 tweets from 1 authors · @random_walker

Biology

New Framework Suggests Camarasaurus Had a Beak-Like Tissue Over Its Teeth

Researchers built a quantitative and qualitative framework linking rostral neurovascular foramina patterns in living archosaurs (birds, crocodylians), lepidosaurs, and turtles to specific perioral soft tissues—carnous lips, corneous beaks, or lipless conditions. Applying this framework to a Camarasaurus specimen, they found its neurovascular signature matches modern birds rather than lipped reptiles, suggesting the sauropod bore a keratinous covering over its complete dentition, a novel condition they term an 'odontoscutum.' The authors propose this keratinized facial tissue may be ancestral to sauropodomorphs, later expanding to cover more of the jaw in derived eusauropods. Since no abstract of the tweets themselves adds independent claims, this summary is grounded in the paper's abstract.

Discussion: 1 tweets from 1 authors · @TomHoltzPaleo

Biology

Endangered Potoroo's Truffle-Only Diet Mapped Across 23 Years of Scat Samples

Researchers used ITS2 metabarcoding on 364 scat samples collected over 23 years to characterize the fungal diet of the long-footed potoroo, an endangered Australian marsupial that feeds almost exclusively on truffle-like ectomycorrhizal fungi. The study found that fungal community composition varies with geography, season, sex, and body mass, suggesting potoroos forage widely at fine spatial scales and disperse fungal spores across microhabitats, with season and its interaction with site being the strongest predictors—raising concerns about how climate change could unevenly affect fungal food resources across the species' two remaining populations. Japanese-language commentary on Twitter highlighted the potoroo as a highly specialized 'fungus-eating' mammal, noting that despite its narrow, truffle-heavy diet it obtains sufficient nutrition via foregut fermentation and passes large numbers of viable spores in its scat, making it an effective spore disperser as well as a dietary specialist.

Discussion: 1 tweets from 1 authors · @Takashirouzu

Computer Science

OpenAgentSafety Framework Tests Real-World AI Agent Risks

The paper introduces OpenAgentSafety, a modular benchmark that evaluates AI agents across eight risk categories using real tools (browsers, code execution, bash shells, messaging platforms) rather than simulated environments, covering over 350 multi-turn, multi-user tasks with both benign and adversarial intents. It combines rule-based checks with LLM-as-judge evaluation, and testing five prominent LLMs revealed unsafe behavior in 51.2% to 72.7% of safety-vulnerable tasks depending on the model, with Claude-Sonnet-3.7 performing best and o3-mini worst. On Twitter, discussion was minimal and largely off-topic: the paper's own author joked that the work has become 'ungoogleable,' apparently overshadowed by a similarly-named or similarly-focused effort from NVIDIA, with no substantive technical critique offered in the visible commentary.

Discussion: 1 tweets from 1 authors · @gneubig

Computer Science

Paper Proposes Economic Frameworks to Manage AI Agents Doing Science

The preprint argues that current multi-agent AI-for-science systems focus too heavily on improving reasoning and hypothesis generation, while neglecting resource management—a critical bottleneck since testing scientific hypotheses is physically and economically costly. The authors propose building 'agentic economies': markets and institutions that let AI agents and human scientists set research priorities, assign credit, track accountability, and guard against misuse, and they discuss the societal implications for equitable distribution of AI-driven discoveries. The paper is conceptual/position-style rather than empirical, sketching infrastructure needs rather than testing a system. The single tweet from one of the authors frames the work as applying an 'economic lens' to identify bottlenecks in AI-driven science and how to address them; there is no substantive external critique or skepticism visible in the discussion so far, which consists mainly of the author's own summary.

Discussion: 1 tweets from 1 authors · @weballergy

Computer Science

Deep Equilibrium Models Replace Deep Networks with Root-Finding

The paper proposes Deep Equilibrium Models (DEQs), which model sequential data by directly solving for the fixed point that hidden layers of deep sequence models tend to converge to, rather than stacking many layers. This is mathematically equivalent to an infinite-depth weight-tied network, but implicit differentiation allows backpropagation through the equilibrium point using only constant memory regardless of effective depth. Applied to transformers and trellis networks on WikiText-103, DEQs match or improve performance while cutting memory use by up to 88%. The Twitter discussion is minimal, with one commenter simply expressing enthusiasm ('this seems interesting') without substantive technical critique.

Discussion: 1 tweets from 1 authors · @ycpAvYHl0gElilF

Biology

PEM-UDE Method Extracts Equations from Noisy Chaotic Neural Data

The paper introduces PEM-UDE, a scientific machine learning method that combines prediction-error feedback with universal differential equations to recover interpretable governing equations from chaotic, noise-corrupted data. The authors validate it on benchmark chaotic systems (Rössler attractor, an electrical circuit) even under heavy noise, then apply it to a simulated population of Izhikevich neurons to derive a reduced-order neural mass model linking single-neuron parameters, sparse connectivity, and macroscopic oscillation/synchrony patterns—predictions they check for consistency against intracranial rat and human cortical recordings. Notably, the authors are explicit that this is an indirect consistency check of trends rather than a direct fit to experimental data. The single tweet highlighting the paper frames it as a meaningful advance for extracting interpretable math from messy chaotic systems, emphasizing its testable predictions for neural networks; no substantive criticism or skepticism appeared in the discussion.

Discussion: 1 tweets from 1 authors · @MillerLabMIT

Other

Japanese Film Studies Paper Analyzes 'Two Anxieties' in Mockumentary Horror

This paper, published in the Japanese journal Eizogaku (Film Studies), examines the mockumentary-horror genre in Japanese cinema, focusing on the works of director Kotaro Terauchi and identifying two distinct forms of anxiety that structure these films. No detailed abstract is available beyond the title and journal listing, so this summary is based primarily on the title and the surrounding Twitter discussion. Terauchi himself shared the piece, expressing surprise and delight at seeing his filmography become the subject of academic analysis, noting how closely the author examined his work.

Discussion: 1 tweets from 1 authors · @terakoh_tera

Mathematics

Game Theory Explains Why Everyone Suddenly Knows Five Katherines

This satirical arXiv paper models baby naming as a competitive game, assuming parents are 'myopic, perfectly knowledgeable agents' who choose names purely to maximize uniqueness. Using this tongue-in-cheek framework, the authors run numerical experiments and even probe large language models to explore naming dynamics, playfully claiming their absurdly simplified model 'perfectly captures the real world.' The paper is a parody of rigorous game-theoretic economics papers, using the format to comedic effect while gesturing at real questions about naming trends and fads. Twitter commentary was minimal but enthusiastic, with one user simply endorsing it as 'completely worth the full ride,' suggesting the humor and cleverness of the piece won over readers who appreciate a well-executed academic parody.

Discussion: 1 tweets from 1 authors · @taracchage

Medicine

Systematic Review Links Heart-Focused Anxiety to Worse Cardiac Outcomes

This systematic review pooled 26 studies (9,628 patients across eight countries) measuring heart-focused anxiety (HFA)—fear of cardiac sensations, hypervigilance, and avoidance of exertion—using validated questionnaires (CAQ or HAF-17) in cardiovascular disease patients. The authors report clinically elevated HFA in 15-49% of patients depending on population and assessment timing, and find HFA independently predicted major adverse cardiac events and mortality post-MI in two prospective cohorts (HR 1.30-1.59). They conclude standard and specialized cardiac rehabilitation, device therapy, and early mobilisation were associated with HFA reductions, though most evidence is observational, and call for routine screening and targeted CBT trials. Twitter discussion around this paper was limited to the first author celebrating publication in a high-impact-factor (IF 10.0) journal, with no substantive critique offered in the visible commentary.

Discussion: 1 tweets from 1 authors · @Dr_NHenryE

Mathematics

Larry Guth Reflects on What Mathematics Means to Him

In this personal essay, mathematician Larry Guth shares his perspective on mathematics—how he views the field, what he values about it, and what it means to him personally, rather than presenting new technical results. The Twitter discussion simply flagged the essay's release with a brief description, without substantive debate or critique visible in the available commentary.

Discussion: 1 tweets from 1 authors · @DynamicsSIAM

Computer Science

Study Finds Most Chain-of-Thought Faithfulness Metrics Perform Near Chance

The paper introduces BonaFide, a benchmark of 3,066 labeled chains-of-thought across 13 tasks and 10 models, built using an automated pipeline that generates ground-truth faithfulness labels rather than relying on proxies like plausibility. Testing existing faithfulness metrics against these ground-truth labels reveals that most perform near chance, show strong prediction biases, and degrade on longer reasoning traces—the best metric only reaches 0.70 AUROC at the CoT level and 0.59 at the step level, with none transferring well across settings and all being computationally expensive. The authors argue this exposes a fundamental gap in how the field currently evaluates whether a model's stated reasoning actually reflects its internal computation. On Twitter, the authors (posting as part of a trio of NeurIPS papers from their lab) framed the core finding succinctly: current faithfulness metrics are "either close to random or prohibitively expensive to run," positioning this as a call to action for more rigorous interpretability tooling rather than a critique of any single prior method.

Discussion: 1 tweets from 1 authors · @megamor2

Physics

Preprint Recasts Periodic Table and Chemical Bonding as 16-Bit Algebraic Vacuum Code

The preprint proposes a speculative algebraic-topological framework in which the field K=ℚ(√2,√3,√5) is treated as a 16-bit 'vacuum register,' with radicals acting as binary opcodes and quadratic subfields forming a 'Fano week' whose discriminants encode a holographic mask. Building on this number-theoretic scaffolding, the authors claim to derive nuclear closure numbers (Z=172), an algebraic UV cutoff at √13, and a reinterpretation of chemical bonding as phase-closed optical interference of confined light rather than electrostatic interaction. The paper explicitly separates rigorous math theorems from more speculative 'ontological' physics postulates, acknowledging the latter's conditional status. The tweets found were self-promotional citations from the authors with no independent critical discussion visible, so no substantive external commentary or skepticism could be captured here.

Discussion: 2 tweets from 1 authors · @JCPEREZCODEX, @JCPEREZCODEX

Computer Science

MeZO Fine-Tunes Huge LLMs Using Only Forward Passes

The paper introduces MeZO, a memory-efficient zeroth-order optimizer that fine-tunes language models using only forward passes (estimating gradients from two forward evaluations) rather than backpropagation, matching inference-level memory use. The authors report it can fine-tune a 30B-parameter model on a single 80GB GPU—versus only 2.7B with standard backprop—while achieving performance comparable to full fine-tuning across classification, multiple-choice, and generation tasks, with up to 12x memory and 2x compute savings; they also give theoretical arguments for why pre-training and prompting let MeZO beat classical zeroth-order slowness predictions. On Twitter, one of the paper's authors noted that community 'research previews' have suggested MeZO generalizes even beyond the settings tested in the paper, but cautioned that such informal reports should be taken with skepticism until verified.

Discussion: 1 tweets from 1 authors · @SadhikaMalladi

Biology

LongevityBench: A New Test of How Well AI Models Understand Aging Biology

No abstract is available for this paper, so this summary is based on discussion only. The authors introduce LongevityBench, an open benchmark designed to evaluate whether frontier large language models (e.g., GPT- and Claude-class systems) genuinely understand aging biology concepts like DNA repair, rather than just producing plausible-sounding text, and they report accompanying language models tailored to this domain. On Twitter, the lead author framed the work as showing that despite the hype around powerful AI models, there remain real gaps in their grasp of aging and longevity biology, with the tweet serving mainly as an announcement of the Cell publication rather than sparking broader debate.

Discussion: 1 tweets from 1 authors · @andrewaiginin

Biology

New Cretaceous Fossil Reveals Semiaquatic Mammal with Continuous Tooth Replacement

The paper describes a newly identified eutriconodontan—an extinct lineage of early mammals—from Early Cretaceous China that shows both semiaquatic adaptations and polyphyodonty, meaning it continuously replaced its teeth throughout life rather than having just two tooth generations like most modern mammals. No abstract was available, so this summary is based on the title and discussion alone. The finding would add to a growing set of Mesozoic mammal fossils (like the beaver-like Castorocauda) showing surprising ecological and dental diversity in early mammal evolution.

Discussion: 1 tweets from 1 authors · @TomHoltzPaleo

Physics

Holographic Dual Proposed for Integrable Field Theories via Topological Strings

The paper proposes a holographic dual for integrable field theories derived from 4d Chern-Simons theory, identifying the dual as a topological string on a generalized Calabi-Yau manifold encoded by a 'planar spectral curve.' The authors describe RG flow as a geometric flow on the moduli space of these spectral curves, valid at all orders in the 't Hooft coupling, and explicitly verify matches with known field theory results up to two loops across several models including the principal chiral model with WZ term and U(N)×U(N) sigma models. The single tweet discussing the paper (in Japanese) expresses admiration for the result but notes it will take considerable time to fully digest the technical content.

Discussion: 1 tweets from 1 authors · @Soliton111

Physics

Paper Argues Von Neumann's Forgotten 1929 Entropy Is the True Thermodynamic One

The paper revisits two entropy definitions proposed by von Neumann: the famous 1927 entropy, which vanishes for all pure states and is now widely used, and a largely forgotten 1929 entropy, which is nonzero for almost all pure states. After tracing the history and mathematical properties of both, the authors argue that the 1927 entropy should be understood as a measure of quantum entanglement rather than thermodynamic entropy, and that the neglected 1929 entropy is actually the correct quantum analogue of thermodynamic entropy.

Discussion: 1 tweets from 1 authors · @tjmlab

Chemistry

Cancer Cells' Own Carbon Monoxide Used to Synthesize Anticancer Drug In Situ

This Nature Communications paper (abstract unavailable) reportedly describes a strategy that exploits endogenous carbon monoxide produced within cancer cells to drive an in-cell chemical synthesis reaction, generating an anticancer drug selectively at the tumor site. The approach falls under "in vivo synthetic chemotherapy," aiming to use a cancer-specific metabolic byproduct as a trigger for localized drug production rather than relying on external activation or delivery of a pre-made drug. RIKEN researchers, who led the work, frame this as a way to improve selectivity and reduce off-target toxicity compared to conventional chemotherapy.

Discussion: 1 tweets from 1 authors · @RIKEN_JP