Rankings

Best LLM Papers

Large language model papers from the last 90 days, ranked by three independent signals: how fast the official implementation is gaining stars on GitHub, how often the paper is cited, and how much attention it drew on Hugging Face. A paper that scores on only one of the three does not make the list. Rebuilt every day.

Ranking window: 2026-07-04 – 2026-10-01 Rankings last rebuilt 2026-10-02

4

LingBot-VLA 2.0: a vision-language-action model pretrained on 60,000 hours spanning 20 robot embodiments

This technical report upgrades LingBot-VLA with a revamped data pipeline of about 60,000 pretraining hours: 50K robot-trajectory hours across 20 embodiments plus 10K egocentric human-video hours. It extends control beyond dual arms to heads, waists, mobile bases and dexterous hands, aiming squarely at the gap between lab robotics and real deployment.

  • 32 stars/7d
  • 15 citations
  • 21 upvotes
  • #3 robotwin-2-0-easy-50-tasks
5

YuE2: one Mixture-of-Transformers that writes a readable score first, then renders full-song audio from it

Symbolic music models plan composition explicitly but never sound like a finished recording, while audio models produce full songs that keep the composition hidden. YuE2 combines both in a single AR-NAR Mixture-of-Transformers that plans symbolically, expands the plan into semantic music tokens, and renders the full song, using two new representation models (MERT2, SheetSage2) to learn from recordings that lack aligned scores. The authors report a SongBench Global Avg of 6.73 on WildSongBench (6.96 at best-of-8), and in expert listening the best-of-8 samples were preferred over Suno v4.5 and roughly tied with Suno v5; the readable score also makes the output editable by external language-model agents.

  • 437 stars/7d
  • 0 citations
  • 216 upvotes
6

Xiaomi Robotics 1: a vision-language-action model pretrained on over 100K hours of real-world trajectories

Xiaomi's robotics team presents a vision-language-action foundation model for mobile manipulation, pretrained on more than 100K hours of real-world UMI trajectories with auto-labeled scene-transition language, then post-trained to match embodiments and human imperative prompts. The paper claims clean scaling with more data and parameters, out-of-the-box performance in unseen environments and efficient fine-tuning for dexterous tasks.

  • 9 stars/7d
  • 15 citations
  • 75 upvotes
  • #3 robocasa
7

WROP: a cognitive-science exam that tests whether video world models understand object permanence — and a 16B model trained to pass it

When an object leaves the frame, does a video generation model remember it exists? The authors built WROP, 150 cognitive-science tasks across six categories rendered in Blender with randomized lighting, speed, and camera angle — a 1.5M-sample training corpus plus a 300-question exam used to test 14 video models. They then trained PWM-WROP, a 16B world model on their native PyTorch stack for AWS Trainium2, releasing data, exam, scores, and weights. In a blind pairwise Elo study, it ranked first among continuation models and effectively tied the best reference-to-video models.

  • 409 stars/7d
  • 0 citations
  • 193 upvotes
8

RoboDojo: one benchmark that puts generalist robot policies through 42 simulation and 18 real-world tasks

RoboDojo is a unified sim-and-real benchmark for generalist robot manipulation policies, with 42 simulation tasks and 18 real-world tasks probing generalization, memory, precision and long-horizon control. It ships parallel Isaac Sim evaluation plus a cloud-accessible real-eval system and a leaderboard the authors populated with 30 policies.

  • 35 stars/7d
  • 12 citations
  • 17 upvotes
9

Frontis-MA1: an open 35B meta-evolution agent for recursive self-improvement in ML engineering

Frontis-MA1 ships with OpenMLE, an open full-stack system (task gyms, RL operator learning, evolutionary search) for studying recursive self-improvement. Its 35B meta-evolution agent, post-trained around Draft, Improve, Debug and Crossover program-evolution operators, claims MLE-Bench Lite results the authors say exceed GPT-5.5 plus Codex, with components transferring to held-out NatureBench Lite.

  • 26 stars/7d
  • 3 citations
  • 186 upvotes
13

Spark-to-Paper: from research idea to submission-ready paper, as thirteen composable skills

Spark-to-Paper implements end-to-end paper generation as thirteen composable skills inside an existing coding assistant, with no separate agent platform: it retrieves literature, designs and runs experiments, revises claims to match evidence and produces figures, keeping model judgment separate from mechanically checkable operations.

  • 73 stars/7d
  • 0 citations
  • 291 upvotes
14

TurboVLA: running vision-language-action policies at ~32 Hz on a consumer GPU, no LLM in the loop

Most VLA pipelines route visual observations through a large language model before decoding actions, which is heavy at inference time. TurboVLA instead encodes vision and language independently, exchanges information through a lightweight bidirectional interaction module, and decodes continuous action chunks directly — a 0.2B-parameter policy the authors report reaching 97.7% average success on LIBERO in 31.2 ms with under 1 GB of VRAM on an RTX 4090.

  • 86 stars/7d
  • 0 citations
  • 140 upvotes
15

LimiX-2: a tabular foundation model that learns data-generating mechanisms instead of just predicting targets

LimiX-2 is a tabular foundation model built on Contextual Mechanism Networks: instead of the usual in-context objective of predicting y given x and context tables, it learns the joint structure p(x, y | context) of how the data was generated. It is pretrained with context-conditional masked modeling on synthetic datasets built from structural causal models spanning varied graph structures and mechanisms. The authors report state-of-the-art results against both dataset-specific models and other tabular foundation models on TabArena, TALENT, and BCCO, and show the mechanism-oriented training has a side effect: the model's feature attention encodes direct causal relationships, allowing it to recover causal skeletons from data.

  • 57 stars/7d
  • 0 citations
  • 188 upvotes

How this list is ranked

Large language model papers from the last 90 days, ranked by three independent signals: how fast the official implementation is gaining stars on GitHub, how often the paper is cited, and how much attention it drew on Hugging Face. A paper that scores on only one of the three does not make the list. Rebuilt every day.

A dash means the signal was unavailable, not zero.