Which Local LLM Performs Autoresearch the Best? (Ornith, Qwen, Nemotron, Muse Glimmer)



Four local models, one DGX Spark, six hours each: which one can run the strongest autonomous research loop and actually improve a training recipe on its own?

🧠 Level up your agents with Agent Wikis — the knowledge bases I use every day to research and build these videos (Hermes Agent, HyperFrames, local AI, and more). Standard wikis are free; Pro ($9.99/mo) adds XL wikis plus custom skills my agents and I have developed over months of working together. Sign up now and you’re locked in at that price for life — it’s going up when agent profiles and workflow templates land: https://agentwikis.com/

Sign up for my FREE weekly newsletter, where I spill my unfiltered thoughts on the latest AI news, cool research, and projects I’m building: https://www.onchainaigarage.com/

A different kind of model test — nothing visual this time. Using Karpathy’s AutoResearch repo (the project that kick-started this channel), four local models each get the same contract: one editable train.py, a fixed compute window, and one metric — validation bits per byte — as they form hypotheses, tweak the recipe, train, evaluate, and keep or discard, fully autonomously on my DGX Spark. The contenders: Ornith, Meta’s Muse Glimmer 30B, NVIDIA’s Nemotron 3.5 Lightning, and Qwen 3.8 27B, with my orchestrator agent running the whole tournament (including overnight). Every run has its own personality — one thinks too much, one won’t stop stopping, one rips through experiments — and the final leaderboard came down to a sliver.

Resources:
🔗 AutoResearch — Karpathy: https://github.com/karpathy/autoresearch

Timestamps:
0:00 – A different kind of model test (no visuals)
2:10 – What AutoResearch is + why it matters to me
4:02 – The rules: one file, one metric, six hours
4:26 – The four local contenders
8:14 – Run 1: Ornith
14:53 – Run 2: Nemotron Lightning (overnight)
19:44 – Run 3: Muse Glimmer
25:31 – Run 4: Qwen 3.8 27B
28:37 – The final leaderboard + takeaways

#AutoResearch #LocalLLM #DGXSpark #Karpathy #Qwen #Nemotron #ModelTesting #LocalAI #MLEngineering

source

Categories:

Leave a Reply

Your email address will not be published. Required fields are marked *

Related Posts :-