Post

Log inSign up

Post

Ricardo Olmedo on X: "We fine-tuned Alec Radford’s 1930 vintage LLM to solve SWE-bench issues. After just ‼️250‼️ training examples, the model solves its first issue, a simple patch to the xarray library. 🧵👇"

  • user avatar
    Ricardo Olmedo
    @rdolmedo_
    We fine-tuned Alec Radford’s 1930 vintage LLM to solve SWE-bench issues. After just ‼️250‼️ training examples, the model solves its first issue, a simple patch to the xarray library. 🧵👇
    7:54 PM · May 2, 2026302.7KViews
  • user avatar
    Ricardo Olmedo
    @rdolmedo_
    May 2
    We scale fine-tuning to ~75k training examples, or 1B tokens. This takes the base model from 4% pass@100 on HumanEval to 4.5% ‼️pass@1‼️ on SWE-bench, a much more challenging agentic benchmark. Having pre-trained only on pre-1931 data!
    user avatar
    Ricardo Olmedo
    @rdolmedo_
    May 2
    Surely a model pre-trained on the web would fare much better? Yes, and no. We also fine-tune their web-retrained model, and observe a modest +1% solve-rate on SWE-bench, achieving 5.7% pass@1 compared to 4.5% Surprisingly little seems to be lost by throwing away the internet.
    user avatar
    Ricardo Olmedo
    @rdolmedo_
    May 2
    What holds the 1930 model back is that it is severely undertrained (only 260B tokens), rather than its pre-training data. I’m excited to see further development of vintage models. Which capabilities does web pre-training provide that are not easily recoverable via post-training?
    user avatar
    Ricardo Olmedo
    @rdolmedo_
    May 2
    Check out the SWE-bench fix discovered after just 250 training examples. The fix itself is very simple, but the trajectory demonstrates the kind of agentic reasoning we’ve come to expect from modern models. 🔗 ricardodominguez.github.io/blogs/pydata__…
    user avatar
    Ricardo Olmedo
    @rdolmedo_
    May 2
    Do you have compute to spare? We’d love to see the full scaling curves comparing the 1930 and web-pretrained models as post-training is scaled up. Models & training data 🔗 huggingface.co/collections/ri… GitHub repo 🔗 github.com/RicardoDomingu…
    1930 Coder - a ricdomolm Collection
    From huggingface.co
    user avatar
    Ricardo Olmedo
    @rdolmedo_
    May 2
    Curious about how pre-training data affects post-SFT performance? Check out Section 4.1 of arxiv.org/pdf/2407.07890
    user avatar
    Ricardo Olmedo
    @rdolmedo_
    May 2
    @status_effects @DavidDuvenaud @AlecRad @jyangballin @OfirPress
  • user avatar
    buge4
    @wanon77789
    May 3
    This shit goes crazy in claude code 😭

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email

Relevant people

Avatar
Ricardo Olmedo@rdolmedo_Follow
PhD student @MPI_IS, working with Moritz Hardt and Bernhard Schölkopf | Currently visiting @Stanford

Trending now

Terms·Privacy·Cookies·Accessibility·US TIDA·Ads Info·© 2026 X Corp.