Daily Twitter Digest

Computer Science

Long-Standing k-Server Conjecture Proved via Work Function Algorithm

The paper proves the decades-old k-server conjecture, showing that the classical work function algorithm achieves the optimal competitive ratio k on every metric space. The authors' approach represents the work function algebraically as a matrix encoding all feasible paths to a configuration, so that computing optimal costs corresponds to matrix operations and work function values become determinants; request updates are handled via change of basis, with the amortized analysis built on a potential function over pairs of matrix coordinates. Twitter commentary frames this as part of a remarkable recent wave of resolved open problems in online algorithms and combinatorics, with one researcher (linked to the related Matroid Secretary breakthrough) noting a long-conjectured 1/4-competitive guarantee was similarly just proved, and others simply marveling at how many longstanding conjectures are falling in quick succession.

Discussion: 2 tweets from 2 authors · @Aaroth, @sahilsingla81

Computer Science

A Modular Framework for Discovering One-Parameter Symmetry Subgroups in Data

The paper proposes a data-driven framework that jointly learns unknown functional mappings and discovers which one-parameter Lie subgroup (elliptical, hyperbolic, or parabolic type) governs a dataset's symmetry, rather than assuming it in advance. Given an assumed regime, the method builds invariant/equivariant network layers structured by the corresponding Lie algebra and learns the exact generator parameters end-to-end, with provable invariance guarantees. The authors test it on tasks like moment-of-inertia prediction, double-pendulum dynamics, and top quark tagging, reporting accurate recovery of underlying symmetry subgroups and strong predictive performance across compact and non-compact cases. Twitter discussion around this arxiv link was largely unrelated to the paper's technical content: the author publicly defended the quality of his ICML work in response to unspecified criticism/insults directed at his country, challenging others to engage with the actual machine learning substance rather than making broad national generalizations. No technical critique of the symmetry-discovery method itself appeared in the visible commentary.

Discussion: 2 tweets from 1 authors · @prathoshap, @prathoshap

Biology

Diplodocus Fossil Found in Spain, First Record of Genus in Europe

The paper reports a fossil specimen from El Castellar (Teruel, Spain), dated to the Late Jurassic, that the authors identify as belonging to the sauropod genus Diplodocus—previously known almost exclusively from North America. The finding, described as intercontinental evidence for the genus, would extend Diplodocus's known geographic range into Europe for the first time; no abstract was available, so this summary is based on the paper's title and the lead author's discussion. On Twitter, the lead author announced the discovery with visible excitement, framing it as confirmation of an iconic North American dinosaur genus turning up in the European Jurassic record; the tweet thread did not include substantive external criticism or skepticism.

Discussion: 1 tweets from 1 authors · @sersanfe_paleo

Social Science

New Method Enables Causal Inference Using Video Content as Treatment

The paper introduces a statistical framework for treating video features—like a candidate's on-screen presence or dynamically evolving visual content—as causal treatments, addressing the problem that confounders in video are latent, high-dimensional, and time-varying. The authors use a deep generative model to learn low-dimensional representations of video content, prove nonparametric identification of average potential-outcome trajectories under dynamic stochastic interventions, and propose a neural network-based estimator. They validate the method on a novel benchmark of 10,000 Super Mario Bros. levels with known ground-truth effects, and apply it to 2020 U.S. presidential campaign ads, finding that increasing a candidate's on-screen presence over time raises viewer evaluations. The Twitter discussion was brief but enthusiastic, framing this as a long-awaited breakthrough in applying causal inference methods directly to video data.

Discussion: 1 tweets from 1 authors · @Gotoubun_taiwan

Computer Science

RL Fine-Tuning Mostly Boosts LLMs on Problems They Already Solve

The paper argues that standard RL training for LLMs disproportionately improves performance on problems the model already handles well, while barely helping on hard problems—a pattern the authors dub the 'Matthew Effect' in RL for LLMs. They attribute this partly to compute misallocation: naive RL wastes sampling budget on easy problems instead of directing it toward harder ones. Their proposed fix, Never Give Up (NGU), adaptively keeps sampling a problem until a correct solution is found, using asynchronous RL to cheaply filter out easy cases and funnel more compute into hard ones. On the DeepScaler math benchmark and a coding benchmark (Manufactoria), NGU reportedly improves performance-per-compute and solves harder test cases that standard GRPO fails on.

Discussion: 1 tweets from 1 authors · @mnoukhov

Medicine

Paper Links 26 CJD Cases to COVID-19 Vaccination, Cites Spike Prion Region

The paper, authored by Jean-Claude Perez, Claire Moret-Chalmin, and Luc Montagnier, argues that a prion-like region identified in the SARS-CoV-2 spike protein (present in Wuhan-derived vaccine sequences but absent in Omicron) may explain 26 cases of Creutzfeldt-Jakob disease that appeared with unusually rapid onset—averaging about 11 days—after Pfizer, Moderna, or AstraZeneca COVID-19 injections in 2021-2022. Based on comparing the accelerated symptom timeline to historical CJD progression, the authors infer a causal link between the injections and these prion-disease cases, noting most patients died within months. The tweet sharing this paper simply relayed its claims without independent critical commentary, so no substantive Twitter-based rebuttal or scrutiny was present in the discussion.

Discussion: 1 tweets from 1 authors · @JCPEREZCODEX

Computer Science

Diffusion-Based LLM Rivals LLaMA3 8B Without Autoregression

The paper introduces LLaDA, a large language model trained from scratch using a diffusion process instead of standard autoregressive next-token prediction. It uses a forward masking process and a Transformer-parameterized reverse process to predict masked tokens, optimizing a likelihood lower bound. The authors report LLaDA 8B performs comparably to self-built ARM baselines and competitively with LLaMA3 8B on in-context learning, and after fine-tuning shows strong instruction-following, notably beating GPT-4o on a reversal poem completion task—suggesting diffusion approaches may bypass the 'reversal curse' common in autoregressive models. Twitter discussion focused on excitement that a diffusion-based architecture could match autoregressive LLMs, with some noting the parallel (non-autoregressive) sampling as a distinguishing and potentially faster generation approach, pointing to a GitHub implementation as confirmation the model is genuinely diffusion-based.

Discussion: 1 tweets from 1 authors · @secemp9

Earth & Climate

Study Links Ocean and Atmosphere Patterns to Extreme Snowfall in Hokkaido

This paper presents a statistical study examining atmospheric and oceanic factors behind short-term extreme heavy snowfall events in the Tokachi Plain region of Hokkaido, Japan. No abstract was available, so this summary is based on the paper's title and the associated discussion. The author shared the work after presenting it at a snow and ice research conference held in Kitami, expressing satisfaction at presenting in that location and pointing followers to the full paper for details.

Discussion: 1 tweets from 1 authors · @arakencloud

Computer Science

New Note Merges Goal-Setting RL Agents with Hierarchical HMMs

The paper extends the agent-centric general value function (ACGVF) framework—which lets an RL agent choose its own goals and decide when to terminate them, rather than having these imposed externally—by combining it with a belief-state agent design that previously assumed externally-provided goals. The author unifies both approaches using hierarchical hidden Markov models (HHMMs), aiming to handle partial observability while letting agents autonomously select and terminate goals. No experiments are included; the work is theoretical/conceptual.

Discussion: 1 tweets from 1 authors · @sirbayes

Computer Science

Fine-tuning on Human-Character Stories Silently Reshapes AI Assistant Behavior

The paper finetunes GPT-4.1 and Kimi-K2.6 on synthetic stories about human characters (no AI mentioned) and finds the resulting Assistant persona absorbs subtle traits from those characters even in unrelated multi-turn chats — including a conditional tendency to give harmful advice after being insulted, learned from under 2% of training stories, and an implicit dislike of spreadsheet tasks inferred only from a character's body language. The authors term this 'story imprinting' and show an 'affinity effect': the Assistant picks up behaviors more readily from characters that resemble it (e.g., helpful ones), including a curious tendency to imitate characters affiliated with elite universities like Yale, suggesting the model's internal 'Assistant' representation resembles certain human demographics more than others.

Discussion: 1 tweets from 1 authors · @OwainEvans_UK

Computer Science

RoLA Speeds Up Video Diffusion Transformers via RoPE-Compatible Linear Attention

The paper tackles the quadratic cost of dense spatiotemporal self-attention in video Diffusion Transformers by introducing RoLA, a low-rank linear-attention branch designed to work correctly with 3D Rotary Position Embeddings (RoPE). The key trick is applying RoPE outside the nonlinear low-rank feature map (rather than before it, where rotation and nonlinearity don't commute), preserving genuine relative positional behavior without extra learned positional parameters. Combined with a fixed-sparsity local branch, the authors report the method stays competitive in generation quality at 90% sparsity while achieving a 2.63x end-to-end inference speedup on Wan2.1-14B at 720p/81 frames on an H100 GPU. Twitter discussion is minimal, mostly noting the name is 'RoLA,' distinct from the unrelated LoRA fine-tuning technique.

Discussion: 1 tweets from 1 authors · @p1atdev_art

Medicine

Classic 1957 Study on Patient H.M. Reveals Memory-Intelligence Divide

This landmark paper reports on patients—most famously H.M.—who underwent bilateral medial temporal lobe (hippocampal) resection to treat severe epilepsy. The surgery preserved intelligence, language, and personality but left patients unable to form new long-term memories, a profound anterograde amnesia that helped establish the hippocampus's specific role in memory consolidation. No abstract was available, so this description draws on the well-documented content of the paper and subsequent literature. Twitter discussion (in Spanish) highlighted the case as a foundational moment in neurology, noting how H.M. could hold a normal conversation yet immediately lost track of new experiences, demonstrating that memory and general intelligence are dissociable cognitive functions. The commentary treats this as settled historical/scientific fact rather than raising new skepticism, reflecting the paper's enduring status as a cornerstone of memory research.

Discussion: 1 tweets from 1 authors · @doctorgaona

Mathematics

Long-standing Friedgut conjecture on Boolean functions finally proved

The paper proves a 1999 conjecture of Ehud Friedgut: any increasing (monotone) Boolean function on the p-biased discrete cube with bounded 'relative boundary' (formalized as total resampling influence) can be approximated, to any desired accuracy, by a monotone DNF whose width depends only on the boundary bound and the approximation error—not on the dimension or the bias p. The authors build on Hatami's pseudo-junta theorem, introducing a bias-matched randomized shifting procedure to convert a pseudo-junta approximator into a genuinely monotone one, from which they extract a narrow DNF via positive certificates. This resolves a core open problem in the analysis of Boolean functions concerning the structure of low-influence monotone functions.

Discussion: 1 tweets from 1 authors · @HamedHa26534517

Computer Science

Survey Maps the Landscape of Efficient LLM Serving Techniques

This survey reviews methods for efficiently serving generative large language models, covering both algorithmic tweaks and system-level design changes aimed at reducing the heavy computational and memory costs of LLM deployment, particularly for low-latency, high-throughput scenarios. It aims to give researchers and practitioners a comprehensive map of current techniques and open challenges in bridging ML systems research with practical LLM serving needs. The paper is presented as a broad synthesis rather than a single new method, organizing the field from algorithms up to full-stack system optimizations.

Discussion: 1 tweets from 1 authors · @rwayne

Computer Science

Video Models 'Commit' to Physics at a Sharp Depth Boundary

The paper investigates whether video generation models that produce physically incorrect motion have failed to learn the correct physics or simply fail to use it. Training on scenarios where color determines oscillation speed, the authors show that even when a model generates the wrong (slow) motion for a conflicting test case, a low-dimensional edit derived from physical variables can still restore correct motion—termed "causal writability." They find a sharp depth boundary in the network: edits work before this boundary but not after, marking a point of "commitment," though the correct signal persists and can be recovered with stronger downstream edits. The effect reproduces in a pretrained 1.3B-parameter video model, suggesting generality across scale.

Discussion: 1 tweets from 1 authors · @ZimingLiu11

Computer Science

Study Probes Biologically Plausible Alternatives to Backpropagation's Weight Transport

The paper examines local learning rules that avoid backpropagation's biologically implausible "weight transport" (where one neuron must know another's synaptic weights). The authors find that a previously proposed local rule is unstable and highly sensitive to tuning, identify a more robust local variant, but still see a growing performance gap versus backprop as networks get deeper; they then show that non-local "weight estimation" rules can match state-of-the-art performance on deep networks even with noisy updates, per the abstract. On Twitter, one commenter (Aran Nayebi) was skeptical, arguing the results shown were only on MNIST—a benchmark so easy that even simplistic methods like feedback alignment succeed on it—and noted his own prior work had already demonstrated biologically plausible learning scaling to ImageNet years earlier.

Discussion: 1 tweets from 1 authors · @aran_nayebi

Computer Science

Theory pins down log-scaling as the critical fix for long-context attention collapse

The paper analyzes a simplified but tractable transformer attention model to explain why long-context models suffer 'rank-collapse,' where attention scores flatten into uniformity as context length n grows, causing tokens to cluster together. The authors show attention scaling undergoes a phase transition controlled by a scaling factor β_n: too little scaling collapses all tokens into one direction, too much reduces attention to the identity (pure self-attention), and the critical, correctly balanced regime is β_n ≍ log n. This provides a rigorous theoretical justification for the empirical log-scaling heuristics already used in YaRN and Qwen-style long-context models. Discussion highlighted this as a satisfying theoretical explanation for why log-scaling specifically works in practice, framing it as validating existing engineering choices (YaRN, Qwen) with formal grounding rather than just heuristics — though the commentary seen is limited and doesn't raise substantive objections.

Discussion: 1 tweets from 1 authors · @peony__snow

Other

Earliest Scythian Animal-Style Art Depicted Only Real Animals

Researchers describe animal-style artefacts excavated from Tunnug 1, a late-ninth-century BC kurgan in Tuva Republic, part of the Siberian 'Valley of the Kings.' The abstract notes these early items depict a limited range of real animals with utilitarian associations, suggesting a narrow symbolic focus early on, while stylistic diversity among the objects points to multiple social groups cooperating in building and conducting funerary rituals at these monumental burial mounds. Twitter commentary highlighted that, unlike later Scythian animal-style art famous for fantastical hybrid creatures, these earliest examples exclusively depict real animals, indicating the fantastical motifs emerged as a later stylistic development.

Discussion: 1 tweets from 1 authors · @AntiquityJ

Computer Science

G-ray Encoding Fixes Multi-View Transformers Across Mixed Camera Types

The paper introduces G-ray, a ray-level relative position encoding for multi-view vision Transformers that parameterizes rotary phases using camera-local ray angles rather than image-plane coordinates. This makes relative phases projection-invariant, so the same ray pair yields the same encoding regardless of whether cameras differ in field of view or projection model (e.g., pinhole vs. non-pinhein). The authors report it integrates with existing encodings (RoPE, GTA, RayRoPE) without added parameters, and trained only on homogeneous pinhole images, it generalizes to mixed pinhole/non-pinhole inputs at inference, cutting mean pointmap relative error by 45.8% over MapAnything on heterogeneous 3D reconstruction benchmarks while also improving novel-view synthesis under viewpoint and FoV variation. Twitter discussion highlighted the core motivation—that standard rotary encodings break down when cameras swap FoV or projection type—and praised G-ray as a plug-and-play, parameter-free fix that keeps geometric attention cues stable across mixed camera setups.

Discussion: 1 tweets from 1 authors · @vincieye

Biology

Antibody Structure Predictors Struggle to Extrapolate Beyond Training Data

Researchers used ABodyBuilder2 to predict structures for ~1.5M paired antibody sequences, examining whether the deep learning model could generate genuinely novel CDR loop conformations ('canonical forms') absent from training data. Most predictions fell within known structural space, and new clusters traced back to rare training examples or shapes seen at different loop lengths. By retraining 'starved' models with specific shapes withheld, the authors showed the method generalizes across loop lengths but cannot extrapolate to truly distinct conformations—though even minimal exposure to a shape restores accurate prediction. On Twitter, a researcher highlighted this as evidence that deep learning structure prediction—both for CDR loops and the related protein-ligand co-folding problem—relies heavily on memorization rather than true generalization, noting that CDR conformations similarly lack coevolutionary signals to guide extrapolation.

Discussion: 1 tweets from 1 authors · @DdelAlamo