Daily Twitter Digest

Computer Science

TailRL Trains Policies to Chase Rare High-Reward Rollouts, Not Just the Mean

The paper argues that optimizing average reward in RL can obscure differences between policies that share the same mean but differ in their likelihood of producing rare, high-reward outcomes—a gap that matters as sampling scales at training and inference time. The authors propose Tail-Likelihood RL (TailRL), which maximizes the log-probability of exceeding a randomly chosen reward threshold, effectively turning reward into a family of binary success events across thresholds. This requires only a small change to the advantage function, and the gradient can be interpreted as a mixture of Best-of-k gradients, letting models benefit more from inference-time sampling without fixing k in advance. Experiments span object localization, maze navigation, GUI grounding, and code optimization, showing TailRL avoids suboptimal solutions by exploiting rare high-reward samples. On Twitter, the authors and collaborators framed TailRL as a drop-in modification to existing RL pipelines that aligns training more closely with Best-of-k style inference, generating interest for its simplicity and broad applicability across domains rather than sparking notable public pushback.

Discussion: 2 tweets from 2 authors · @rsalakhu, @di_zhang_fdu

Physics

Why Quantum Mechanics Really Lives in Rigged Hilbert Space, Not Just Hilbert Space

The paper argues that when continuous spectra are involved, the proper mathematical home for quantum mechanics is the rigged Hilbert space rather than the ordinary Hilbert space, since only the former fully realizes Dirac's bra-ket formalism. Using a simple, exactly solvable example, the authors give a constructive, pedagogical derivation of this structure and discuss its physical significance, building on their earlier work. Twitter commentary was limited to a single widely shared reaction lamenting that this mathematical subtlety is rarely even mentioned as a footnote in standard quantum mechanics textbooks, implying textbooks give students an incomplete picture of the theory's rigorous foundations.

Discussion: 1 tweets from 1 authors · @mbatubayindirli

Computer Science

2018 Report Warned of AI-Enabled Hacking, Drones, and Disinformation

This multi-author report surveys how AI could reshape security threats across digital, physical, and political domains, covering scenarios like automated cyberattacks, weaponized autonomous drones, and AI-driven disinformation. The authors propose four high-level recommendations for researchers and policymakers and identify research areas that could strengthen defenses, while leaving open the long-term balance between attackers and defenders. Revisiting the report, one of its co-authors noted that scenarios once dismissed as sci-fi speculation—like automated spear phishing and lethal autonomous drones—are now viewed as more plausible or already emerging, underscoring how quickly perceptions of AI risk have shifted.

Discussion: 1 tweets from 1 authors · @Miles_Brundage

Computer Science

Repeating the Prompt Twice Boosts Non-Reasoning LLM Accuracy

The paper claims that simply duplicating the input prompt—feeding the model the same text twice before generation—improves performance on non-reasoning tasks across Gemini, GPT, Claude, and Deepseek models, without increasing output tokens or latency. The mechanism resembles giving earlier tokens a form of bidirectional-like context: causal attention means the second copy's early tokens can now 'see' information from later in the first copy, something a single forward pass can't provide. Twitter commentary connected this to a known trick used by AI-text detector Pangram, noting it's a workaround for the fact that autoregressive LLMs lack true bidirectional attention, and quipped that this hints at why text-diffusion-style approaches (which aren't strictly left-to-right) might be appealing.

Discussion: 1 tweets from 1 authors · @_chenglou

Earth & Climate

Energy Limits on Soil Microbes May Cap Carbon Storage Capacity

Based on Twitter discussion (no abstract available), the paper by Chaoqun Wang and colleagues in Nature Geoscience argues that soil microorganisms—bacteria and fungi—are fundamentally constrained by limited available energy, and that this energy limitation, rather than substrate supply alone, restricts how much carbon soils can accrue and retain. The authors reportedly compare growth versus maintenance energy budgets and carbon-use efficiencies across microbial groups, then upscale these cell-level findings to global carbon cycling estimates. This reframes soil carbon storage as fundamentally an energetic bottleneck problem for microbial communities rather than purely a matter of organic matter input.

Discussion: 2 tweets from 1 authors · @ykuzyakov, @ykuzyakov

Physics

Paper Finds an Error in Feynman's Gravitation Lectures—But It Still Works

The paper revisits Feynman's derivation of gravity as a massless spin-2 field in flat spacetime, building the theory order-by-order via self-consistency conditions on the Lagrangian. The authors present detailed third- and fourth-order Lagrangian density calculations and show that Feynman's published third-order expression does not actually satisfy his own self-consistency condition—yet it still reproduces the correct perihelion shift of Mercury, one of general relativity's classic tests. This suggests an inconsistency in the historical derivation that nonetheless happened not to affect its most famous physical prediction. Twitter discussion was minimal, largely just flagging the paper's existence with a pointed emoji, though the finding itself is notable for physics historians and those interested in the field-theoretic derivation of GR from flat-spacetime spin-2 field theory.

Discussion: 1 tweets from 1 authors · @pascalkwanten

Medicine

Self-Supervised Diffusion Bridge Reconstructs MRI Without Clean References

The paper introduces SelfDB, a self-supervised diffusion bridge model for MRI reconstruction that trains without high-quality reference images. It works by further sub-sampling available noisy measurements twice and training a network to reverse this degradation, using only the noisy measurements themselves as targets. The authors report SelfDB outperforms standard denoising diffusion models on compressed sensing MRI tasks. The Twitter discussion doesn't engage with the paper's technical claims directly; instead, one user mentions using an AI agent (Astra) with compute credits to explore and attempt to extend this diffusion research, framing it as an experiment in AI-assisted research rather than a critique or validation of the method itself.

Discussion: 1 tweets from 1 authors · @hrygao

Medicine

New Evidence-Based Guidelines for Adult-Onset IgA Vasculitis Published

This Nature Reviews Rheumatology paper presents evidence-based guidelines for diagnosing and managing adult-onset IgA vasculitis (IgAV), a systemic small-vessel vasculitis. No abstract was available, so this summary is based on discussion only. A tweet highlighted that the guidelines are open access and include a figure comparing the histological and clinical similarities and differences between IgA nephropathy and IgAV with nephritis, suggesting the paper addresses diagnostic overlap between these related kidney conditions.

Discussion: 1 tweets from 1 authors · @rheum_showa

Engineering

Graph-as-Policy Harness Aims to Boost Reliability in Variational Robot Automation

The paper introduces Graph-as-Policy (GaP), a multi-agent coding harness that builds directed computation graphs combining perception, planning, and control nodes drawn from a Modular Open Robot Skill Library. GaP rehearses tasks in an internal simulator across parallel graph variants, iteratively refining structure and parameters to improve success rate and throughput on "Variational Automation" tasks—those with more object and pose variability than fixed automation. Across 8 new benchmarks (4 simulated, 4 real-world), the authors report GaP significantly outperforms baseline approaches, aiming to close the reliability gap that model-free policies struggle with in commercial settings. Twitter commentary was brief and mostly promotional: Ken Goldberg highlighted GaP as a newly accepted CoRL paper offering harness-design improvements—specifically adding graphical structure—that could enhance performance on other systems like Astra, framing it as part of a fast-moving wave of agentic robotics benchmarking rather than raising direct criticism.

Discussion: 1 tweets from 1 authors · @Ken_Goldberg

Biology

Perspective Argues ML-Assisted Directed Evolution Misses Real-World Cost Constraints

This perspective piece by Bruce Wittmann argues that despite five years of machine learning advances in protein engineering, machine-learning-assisted directed evolution (MLDE) has largely failed to transform practical directed evolution work. The author contends this stems from a mismatch in goals: MLDE researchers optimize for finding the single best protein, while practitioners need a sufficient protein under real time and resource limits—critically, most MLDE methods ignore DNA synthesis costs, undermining their practical value regardless of model quality. The piece closes by highlighting recent exceptions that better integrate cost considerations, arguing the two goals can be reconciled. Twitter commentary from Anthony Gitter praised the perspective as "phenomenal," specifically highlighting the point that ignoring library construction and cost considerations represents a significant gap between researchers who focus on model training/benchmarking and those actually applying these methods in practice.

Discussion: 1 tweets from 1 authors · @anthonygitter

Computer Science

HELIX Proposes Co-Evolving Agent Harnesses and Models for Self-Improvement

The paper introduces HELIX, a framework that treats an AI agent's runtime harness—the code mediating context, tools, control flow, and stopping conditions—as a first-class, evolvable component alongside the underlying model. HELIX decomposes agent systems into typed ports, reusable atoms, recipes, and runtime policies, making harness modifications explicit, auditable, and source-traceable, while logging trajectories and outcomes for downstream model training. In a single evolution round on code repair (SWE-bench-style evaluation), a 65-candidate harness portfolio improved task coverage by 4.0% over a baseline harness, and the full portfolio exposed up to 58.0% more verified coverage via complementary sibling behaviors, yielding 438 verified training records (SFT, critic, filter, preference) from a 200-slot sibling slice. The authors frame this as a feedback loop: harness evolution expands current capability and produces learning signal for model updates, which in turn motivate further harness evolution. The tweet surfacing this paper was part of a curated reading list shared ahead of a technical meetup on 'agent harness engineering and self-evolving agents,' offering little independent commentary or critique—so the discussion here reflects curation/interest signal rather than substantive debate about the paper's methodology or results.

Discussion: 1 tweets from 1 authors · @ishandutta0098

Computer Science

Study Finds AI-Text Detectors Flag Honest Editing More Than Evasive Rewrites

The paper reports a controlled study of published English abstracts (2013–2015 vs 2023–2025) showing that commercial AI detectors cannot reliably separate light AI-assisted editing from full LLM drafting. Guideline-compliant 'refine abstract only' edits were flagged as AI-generated 38–80% of the time, while unmodified recent human-written abstracts were flagged 9–15%, with non-STEM texts flagged far more often than STEM ones. After running AI text through a 'humanizer' tool, detection rates dropped below 4%, meaning detectors punished honest AI-assisted editing more harshly than deliberate evasion — leading the authors to argue detector scores should not be used as standalone misconduct evidence. The Twitter discussion (in Turkish) centered on sharing the paper as evidence of weaknesses in software claiming to detect AI-written text, with the original poster highlighting the finding as a critique of such detection tools' reliability.

Discussion: 1 tweets from 1 authors · @say_cem

Mathematics

A Graduate-Level Graph Theory Textbook Lands on arXiv

This is a 454-page graduate-level introduction to graph theory, covering simple graphs, multigraphs and their directed analogues, plus restrictive classes like tournaments, trees and arborescences. It presents core results including Eulerian circuits, Hamiltonian cycles, spanning trees, the matrix-tree and BEST theorems, proper colorings, Turán's theorem, bipartite matching, and the Menger and Gallai–Milgram theorems, alongside a network-flows treatment used to prove Hall's marriage theorem, with roughly a hundred exercises (no solutions). The text is framed as material for a quarter-long course rather than a novel research contribution. Twitter commentary was brief and purely appreciative, framing it as a useful free resource for learning graph theory and network science rather than raising any substantive critique.

Discussion: 1 tweets from 1 authors · @KirkDBorne

Computer Science

BLASt3R Unifies VSLAM and Offline SfM in One Bundle Adjustment Framework

The paper presents BLASt3R, a regularized bundle adjustment (BA) system that combines a fast multi-view matcher with monocular depth priors for initialization and regularization, addressing the cost bottleneck of dense correspondence estimation in hybrid Structure-from-Motion pipelines. Unlike prior systems, it uses a single unified optimization framework and shared hyperparameters to handle both online Visual SLAM and offline reconstruction from unordered image collections. The authors report improved speed/accuracy tradeoffs over feed-forward, traditional, and hybrid baselines, and claim their uncalibrated VSLAM variant even outperforms previous calibrated methods. Twitter discussion was limited to a brief technical note characterizing BLASt3R as a combination of MUSt3R and MASt3R matching techniques with depth priors and bundle adjustment, without substantive critique or debate.

Discussion: 1 tweets from 1 authors · @zhenjun_zhao

Biology

Prefrontal High-Frequency Oscillation Chains Mark Distinct REM Memory Replay Pattern

Using high-density tetrode recordings in rodent hippocampus (CA1) and prefrontal cortex, the authors identify chains of prefrontal high-frequency oscillations (HFOs) during REM sleep that are phase-locked to theta rhythms and coincide with increased CA1-PFC theta coherence. These HFO chains organize sparse, temporally extended reactivation of PFC ensembles against a backdrop of local neural suppression—contrasting with the widespread, bursty reactivation seen during NREM sharp-wave ripples—and are linked to shifts in CA1 theta-phase preference and firing-rate changes, suggesting REM sleep actively regulates hippocampal excitability. A companion cortical network model incorporating acetylcholine effects reproduces both the sparse REM and dense NREM reactivation patterns, offering a mechanistic account for why the two sleep stages support memory consolidation differently. The tweet from a neuroscience lab highlights the core finding that PFC fast oscillations coupled to theta may support memory replay during REM sleep, framing it as a novel physiological signature of REM-based consolidation; no substantive critical discussion appeared in the available tweets.

Discussion: 1 tweets from 1 authors · @MillerLabMIT

Medicine

SPARTAN Sets Data-Driven Referral Criteria for Suspected Axial Spondyloarthritis

Given diagnostic delays of up to 14 years for axial spondyloarthritis (axSpA), a multidisciplinary SPARTAN group combined systematic review/meta-analysis of 28 clinical, lab, and imaging features with Delphi consensus and discrete choice experiments to build a referral tool. Features like uveitis, elevated ESR/CRP, HLA-B27, inflammatory bowel disease, imaging sacroiliitis, psoriasis, family history, and NSAID response were weighted via likelihood ratios into a scoring system (threshold score ≥3) for referring adults under 45 with chronic back pain to rheumatology, targeting a ≥33% post-referral diagnosis rate. The authors note the strategy still needs prospective validation to confirm it actually reduces diagnostic delay. Twitter discussion consisted mainly of the journal's own announcement, with no substantive independent critique yet visible in the conversation.

Discussion: 1 tweets from 1 authors · @ACR_Journals

Other

Newly Found Manuscript Quadruples Known Proverbs of 17th-Century Jesuit Hermodorus

This philology article edits and traces the textual history of a proverb collection compiled by the Greek Jesuit Hermodorus Rhegius (1579–1655), who worked mainly on Chios and left important evidence of early modern Greek language and culture. Long known only through 27 proverbs cited by Charles Du Cange, plus three more found in a Paris manuscript in the early 20th century, the collection has now been expanded by 90 newly discovered proverbs in a manuscript at the Médiathèque d'Orléans, bringing the total to 120. The authors also reconstruct a small network of early modern French scholars who took a keen interest in Greek proverbs. Twitter discussion was limited to a single note (in Japanese) flagging the publication as a notable rediscovery for historians of the Greek language, with no substantive critical commentary offered.

Discussion: 1 tweets from 1 authors · @C_Pompilius

Medicine

IVUS-Guided Stenting Benefits Vary Sharply by Geography, Meta-Analysis Finds

This meta-analysis of 17 randomized trials (14,033 patients) compared IVUS-guided versus angiography-guided PCI with drug-eluting stents, testing whether benefits differ by geography as a pre-specified primary hypothesis. Overall, IVUS guidance reduced cardiac death, target-vessel MI, target-vessel revascularization, MACE, and stent thrombosis, but significant interaction tests showed these benefits were markedly larger in East Asian trials (e.g., cardiac death RR 0.56) than in non-Asian trials (RR 1.23, not significant). Stent thrombosis reduction was the exception, remaining consistent across regions. The authors speculate that differences in how operators use IVUS data to optimize stent implantation may explain the divergence. Twitter discussion (from a cardiology journal editor) highlighted the study's framing around this Asian vs. non-Asian split in trial outcomes for major adverse cardiac events, though substantive critical debate wasn't present in the available tweets.

Discussion: 1 tweets from 1 authors · @ehj_ed

Mathematics

Closed-Form KL Divergence Found for Cauchy Distributions

The paper derives a closed-form expression for the Kullback-Leibler divergence between Cauchy distributions, resolving a previously open definite integral. The authors show that, unlike most parametric families, this divergence between Cauchy densities is always finite and symmetric. This makes Cauchy distributions notable exceptions to the general asymmetry of KL divergence.

Discussion: 1 tweets from 1 authors · @FrnkNlsn

Computer Science

Text-to-Image Models Show Stereotyped Objects Tied to Demographic Cues

The paper introduces SODA, a framework using automated attribute discovery and three metrics (BDS, CDS, VAC) to audit demographic bias not in human depictions but in generated objects like cars and laptops. Testing 8,000 images across five leading text-to-image models and eight object categories, the authors find neutral prompts default to outputs resembling middle-aged White people, while explicit demographic cues trigger extreme stereotyping—26.6% of object-model-demographic combinations yielded identical attributes across all 20 generated images (e.g., rose gold laptops for women). They also report that prompt-level debiasing reduces disparity between groups but collapses diversity within groups, effectively swapping one stereotype for another. Twitter discussion so far is limited to a conference-poster announcement, offering no substantive critique or debate beyond promoting the ECCV presentation.

Discussion: 1 tweets from 1 authors · @solsol_choi