Daily Twitter Digest

Computer Science

Open Recipe Trains Qwen2.5-1.5B-Level LLM for Under $7K on Consumer GPUs

The authors present an open, fully-reproducible pretraining recipe for a 2B-parameter language model, Puro-2B, trained from scratch on up to 1.4 trillion tokens using FP8 precision on consumer RTX 5090 GPUs. Combining hardware selection, low-precision training, a 'hyperball' optimization approach, curriculum model averaging, and a tailored data recipe, their best model approaches Qwen2.5-1.5B performance at a training cost under $6.9K—versus over $1.5M for comparable models like Llama-3.2-3B. They also derive a cost-scaling law suggesting ~$4.4K suffices to match Qwen2-1.5B, and use their full pipeline access (not just weights) to study how pretraining data curricula affect downstream post-training performance. Data, code, and weights are released under Apache 2.0. Commentary highlighted the detailed technical report and its use of the MuonH optimizer, drawing comparisons to related 'Effective Learning Rate' (ELR) research, with some noting convergent insights across independent efforts on this topic.

Discussion: 2 tweets from 2 authors · @_reachsumit, @zhanpeng_zhou

Mathematics

Trefethen Reflects on the Status of Three Millennium Prize Problems

In this essay, Nick Trefethen offers reflections on the Riemann Hypothesis, P vs. NP, and the solvability of the Navier-Stokes equations, three of the seven Millennium Prize problems, discussing their current status and what progress (or lack thereof) reveals about the nature of these deep open questions. Twitter commentary framed the piece as a meditation on mathematics' unique ability to pose precise problems that can remain unresolved for decades or centuries yet still be unambiguously solvable in principle, with users sharing it as a thoughtful, accessible read rather than raising substantive criticism.

Discussion: 2 tweets from 2 authors · @BahramShakerin, @miniapeur

Computer Science

Prefix Sliding Cuts Memory Cost of Long-Horizon LLM Reasoning

The paper proposes Prefix Sliding, a method for test-time scaling in language models that discards intermediate reasoning tokens except for a fixed prefix (key instructions/tools) and a sliding window of the most recent tokens, capping memory use regardless of reasoning length. The authors report that, without any retraining, this makes existing models 3x faster while preserving performance, and that training with Prefix Sliding via reinforcement learning enables scaling to reasoning traces beyond 100,000 tokens with better results than either summarizing old tokens or using a vanilla sliding window. The core motivation is an empirical observation that most intermediate reasoning tokens quickly lose relevance as reasoning progresses, making full-attention retention wasteful.

Discussion: 1 tweets from 1 authors · @Muennighoff

Mathematics

New Apéry-Type Proofs Bound Irrationality Measures of q-Series Values

The paper constructs a three-parameter family of rational (Apéry-type) approximations to q-hypergeometric series, using them to prove irrationality of specific values (at r=x^{-1}, |x|≥2) of Ramanujan's theta function ψ(r), the odd-divisor generating function Δ(r), and the 4-regular partition generating function B₄(r). The authors further derive explicit upper bounds on the irrationality measures of these values (18/7, 18π²/(7π²−24), and 3 respectively), and show one approximation recovers a known Padé approximation due to Coussement–Smet.

Discussion: 1 tweets from 1 authors · @J_Koizumi_233

Computer Science

FlashNormal Uses Phone Flash/No-Flash Photo Pairs for Detailed Surface Normals

The paper proposes FlashNormal, a diffusion-based method that estimates high-fidelity surface normals from a pair of flash and no-flash smartphone images, exploiting flash-induced shading changes plus a curvature-guided enhancement strategy to recover fine surface detail and reduce shape-reflectance ambiguity — problems that plague single-image estimators, while avoiding the multi-light capture rig required by classical photometric stereo. The authors also introduce EvalFlash, a new 20-object real-world benchmark with ground-truth normals, and report state-of-the-art results versus single-image methods and prior flash/no-flash approaches. Twitter commentary on this thread was limited to a brief share of the abstract with no substantive critique offered yet.

Discussion: 1 tweets from 1 authors · @ssh4net

Social Science

Study Finds Sleep Incentives Boosted Grades and Habits in Students

Based on discussion only (no abstract available): the authors—Marta Serra-Garcia (@Osea82), Silvia Saccardo, and Sally Sadoff, publishing in the Journal of Political Economy—report a field study where incentivizing students to sleep more increased the share of weeknights they got at least 7 hours of sleep by 26% and raised grades by roughly 0.10–0.11 standard deviations, suggesting the intervention also fostered lasting sleep habits. The tweet frames this as evidence that simple incentive schemes can meaningfully improve both sleep behavior and academic performance. Commentary so far is limited to the author's own summary announcing the publication, with no substantive external criticism yet visible in the discussion.

Discussion: 1 tweets from 1 authors · @Osea82

Physics

New Analysis Argues Palomar Plate 'Transients' Show Real Optical Aberrations

The paper revisits the long-debated VASCO catalog of fast, unexplained brightness transients found in 1950s Palomar sky survey photographic plates. The authors argue that these transient images display a coma aberration pattern consistent with genuine off-axis point sources passing through the telescope optics—a signature they say plate defects or emulsion artifacts cannot naturally reproduce. While the study stops short of identifying what physically caused the light, the authors present it as evidence against the artifact hypothesis that has been used to dismiss the transients.

Discussion: 1 tweets from 1 authors · @DrBeaVillarroel

Physics

Claim of Room-Temperature Superconductivity in Fe-Se-H at Normal Pressure

The paper reports EPR measurements of microwave power absorption versus magnetic field (up to 6000 Oe) in a powdered Fe-Se-H compound, made by diffusing hydrogen into an FeSe single crystal (Tc=8K) under thermal conditions. The authors say the shape of these absorption curves indicates superconducting behavior both in the 3.6-25 K range and, notably, at 295 K (room temperature) under normal pressure. Twitter commentary was limited to a single excited reaction highlighting the room-temperature superconductivity claim, using the classic "here it comes" meme reaction, without offering independent scrutiny or corroboration of the result.

Discussion: 1 tweets from 1 authors · @tjmlab

Medicine

Review Highlights Aldosterone Synthase Inhibitors as New CKD Treatment Strategy

This review argues that aldosterone drives CKD progression via blood pressure effects plus renal inflammation and fibrosis, leaving residual risk even after RAS inhibitors, mineralocorticoid receptor antagonists, and SGLT2 inhibitors. The authors summarize evidence that new selective aldosterone synthase inhibitors (baxdrostat, lorundrostat, vicadrostat) lower blood pressure and reduce albuminuria in CKD patients—sometimes more so combined with empagliflozin—while causing modest, generally manageable increases in serum potassium, with SGLT2 co-therapy potentially mitigating hyperkalemia risk. They note several phase 3 trials are underway to determine if these albuminuria benefits translate into improved kidney and cardiovascular outcomes. Twitter discussion of the paper was limited to a single promotional share from the journal's account, with no substantive critical commentary attached.

Discussion: 1 tweets from 1 authors · @CKJsocial

Computer Science

mold Linker Uses Full Parallelism to Slash C++ Link Times

The paper presents mold, a Unix/Linux linker that applies data parallelism throughout the entire linking pipeline rather than in isolated stages. The authors identify architectural bottlenecks in existing linkers—like entangled symbol resolution and archive processing—that limit parallel scaling, and redesign the pipeline to decouple these steps. On large real-world C++ programs, mold links multi-gigabyte debug binaries in seconds (often under one), reporting 2.4–16.1x speedups over lld and up to 112x over GNU ld, with an ablation study showing gains come cumulatively from parallelizing all passes rather than any single trick. Twitter discussion mainly just shared the paper and venue (ASPLOS 2027) with no substantive critique noted.

Discussion: 1 tweets from 1 authors · @matt_dz

Computer Science

Training-Free 'Recirculation' Technique Cuts Perplexity on Gemma3 Models

The paper proposes 'recirculation,' an inference-time architectural tweak for off-the-shelf foundation models that injects activations from a later layer back into an earlier one, creating a form of recurrence without retraining weights. The authors argue this addresses a fundamental limit of feedforward transformers—bounded state updates across depth—letting models better track belief states, distinct from chain-of-thought or looping approaches. On Gemma3 models, an adaptive variant reportedly yields a 23% perplexity reduction and a 21% accuracy gain on GSM8k, with no added generation latency (though prefill becomes serial). A Twitter user independently reproduced the core effect on Gemma 3 1B, confirming that injecting activations from layer 11 into layer 4 produces the claimed initial perplexity drop, lending early external credibility to the method.

Discussion: 1 tweets from 1 authors · @YangWang92

Medicine

European Heart Journal Paper Makes Case for Early LDL-C Lowering

This European Heart Journal article argues for reducing cumulative LDL-cholesterol exposure through earlier initiation of lipid-lowering treatment, based on the concept that lifetime LDL-C burden—not just current levels—drives atherosclerotic cardiovascular risk; no abstract was available, so this summary is based on discussion only. The piece appears to synthesize existing epidemiological and trial evidence to support treating high LDL-C sooner rather than waiting for risk thresholds to be met later in life. Twitter engagement was limited to a promotional share from the journal's editorial account, with no substantive independent commentary or criticism surfacing in the discussion.

Discussion: 1 tweets from 1 authors · @ehj_ed

Earth & Climate

FAO's 2026 Report Reassesses Global Soil Health a Decade After Landmark Study

No abstract is available for this report, so the following is based on discussion only. The FAO's Global Soil Partnership and its technical panel (ITPS) have released a follow-up to their landmark 2015 Status of the World's Soil Resources assessment, evaluating how soil conditions and the threats facing them—such as erosion, nutrient loss, and contamination—have evolved over the past decade. The tweet from FAO Knowledge simply announces the report's release and frames it as a decadal check-in on global soil health, without engagement from independent commentators offering critique or additional context.

Discussion: 1 tweets from 1 authors · @FAOKnowledge

Social Science

Leiden Summer School Releases Comparative Semitics Course Syllabus

No abstract is available for this item, so this summary is based only on the tweet discussion. The linked resource is a syllabus from a course on Comparative Semitics taught at the Leiden Summer School, covering the comparative study of Semitic languages (e.g., Arabic, Hebrew, Aramaic, and related tongues) and their historical relationships. According to the tweet, the instructors cleaned up their teaching materials and released them under a Creative Commons license for public use. Twitter commentary was limited to the announcement itself, framing it as an openly shared educational resource for those interested in Semitic linguistics, with no substantive criticism or debate visible in the discussion.

Discussion: 1 tweets from 1 authors · @PhDniX

Biology

AI Models Match Human Experts at Designing RNA Pseudoknots

Researchers from the Eterna project and Das Lab tested five AI teams against experienced human designers on 57 blind RNA pseudoknot design challenges, evaluating results with chemical mapping, compensatory mutagenesis, and cryo-EM. Generative AI methods matched human performance on most tasks, and successful designs formed well-ordered 3D structures stabilized by tertiary interactions the models never explicitly modeled—suggesting some RNA design problems may be solvable without first cracking full 3D structure prediction. The work also comes with an accompanying commentary on RNA structure prediction from the Mustoe lab. The lead author highlighted the collaboration between Eterna citizen scientists, five AI teams, and cryo-EM structural validation as a key feature, framing the accompanying commentary as 'provocative' regarding the state of RNA structure prediction more broadly.

Discussion: 1 tweets from 1 authors · @RDasLab

Social Science

Review Synthesizes Criteria for Evaluating Qualitative Research Rigor

This review synthesizes published evaluative criteria for judging the rigor of qualitative research, drawing on a systematic search of journal articles and their references to identify well-defined markers of quality across diverse epistemological traditions. The authors conclude that because qualitative research spans multiple paradigms and is not a unified discipline, no single universal checklist of quality criteria is feasible; instead, researchers should evaluate studies using standards appropriate to the theoretical and methodological framework from which they originate. The paper offers recommendations aimed at improving how qualitative work is assessed and conducted. Twitter commentary consisted of a brief share of the citation with minimal added discussion, suggesting interest mainly in circulating the reference for others working in qualitative methods.

Discussion: 1 tweets from 1 authors · @tomokoba10

Computer Science

ExMesh++ Turns Multi-View Photos Directly into Relightable UV-PBR Mesh Assets

ExMesh++ proposes a staged pipeline for converting multi-view images into production-ready 3D mesh assets with proper topology, UV parameterization, and explicit PBR material maps. Rather than jointly optimizing geometry, materials, and lighting (which the authors say causes ambiguous decomposition), the method first refines mesh topology via adaptive vertex splitting/merging while preserving UV consistency, then fixes that mesh-UV "carrier" to optimize PBR maps and environment lighting, including one-bounce diffuse indirect illumination via secondary-ray tracing. The authors report competitive geometry accuracy and strong relighting results, with assets directly usable in standard DCC (digital content creation) software. The one tweet covering this highlighted the practical appeal: unlike typical reconstruction pipelines that yield messy topology, missing UVs, and shadows baked into textures, ExMesh++ promises clean, artist-ready assets with disentangled materials and lighting straight out of multi-view capture — a workflow pain point for graphics practitioners.

Discussion: 1 tweets from 1 authors · @vincieye

Computer Science

Praxist Open-Sources Lineage-Tracking System for Autonomous R&D Agents

Praxist is a system for autonomous R&D agents that builds a typed 'evidence graph' linking experimental artifacts to the findings and design decisions that produced them, letting later attempts inherit validated mechanisms rather than re-discovering the same lessons. On the MLE-bench suite (75 tasks), the authors report Praxist achieving 60 medals (80.0%, 49 gold) versus 55 medals (73.3%, 34 gold) for a Claude Code/Opus 4.8 baseline, at roughly one-twelfth the model spend ($3,054 vs $38,370), with additional case studies in trading, SLAM, tokamak control, and rocket landing. The announcement centers on the code, tech report, and GitHub release now being made public, framed by the authors as moving agentic R&D from benchmark demos toward auditable, reproducible engineering practice; the linked tweet is promotional rather than critical, so independent scrutiny of the cost and medal claims has not yet surfaced.

Discussion: 1 tweets from 1 authors · @Sapient_Int

Computer Science

Study Links LLM Performance Gap to Representational Geometry Degeneration

The paper examines why LLMs perform worse on low-resource languages by analyzing the geometric properties of hidden representations across 30 languages, finding that low-resource languages exhibit "representational degeneration" particularly in final layers, correlating with data availability. The authors test geometric regularization during continued pretraining on 9 base LLMs adapted to 10 African languages, showing it reduces this degeneration and yields modest performance gains, especially on harder tasks. The work suggests targeted geometric interventions could help close the low-resource language performance gap in LLMs. The author announced the paper's acceptance to EMNLP Main on Twitter, framing it as a study of how representation geometry differs across LLM layers and between low- and high-resource languages; the tweet thread offers no independent critique or skepticism from other users.

Discussion: 1 tweets from 1 authors · @francoisrmeyer

Social Science

Neural Similarity May Predict Who Becomes Friends, Study Suggests

No abstract was available for this paper, so this summary is based solely on the Twitter discussion. According to a tweet describing the Nature Human Behaviour study, researchers scanned participants' brains and found that greater similarity in neural activity patterns—particularly within default mode network regions—predicted which individuals would go on to form friendships, suggesting friendship formation may be linked to shared neural response profiles. The commentary highlights this as a striking finding but offers no independent scrutiny of methodology, sample size, or causal direction, so these claims should be treated cautiously pending further detail from the paper itself.

Discussion: 1 tweets from 1 authors · @virginiog