<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>single-file on tomrochette.com</title>
    <link>https://tomrochette.com/tags/single-file/</link>
    <description>Recent content in single-file on tomrochette.com</description>
    <generator>Hugo -- gohugo.io</generator>
    <language>en</language>
    <managingEditor>tom@tomrochette.com (Tom Rochette)</managingEditor>
    <webMaster>tom@tomrochette.com (Tom Rochette)</webMaster>
    <copyright>© 2026 Tom Rochette</copyright>
    <lastBuildDate>Wed, 07 Oct 2026 06:38:49 -0400</lastBuildDate><atom:link href="https://tomrochette.com/tags/single-file/index.xml" rel="self" type="application/rss+xml" />
    
    <item>
      <title>llamafile</title>
      <link>https://tomrochette.com/agents/model-access/llamafile/</link>
      <pubDate>Wed, 07 Oct 2026 00:00:00 +0000</pubDate>
      <author>tom@tomrochette.com (Tom Rochette)</author>
      <guid>https://tomrochette.com/agents/model-access/llamafile/</guid>
      <category>research-note</category><category>agent-curated</category><category>fully-ai-generated</category><category>llm=glm-5.3-flash</category><category>model-access</category><category>local-inference</category><category>single-file</category>
      <description>&lt;p&gt;llamafile is Mozilla&amp;rsquo;s single-file LLM distribution format: it folds llama.cpp and Cosmopolitan Libc into one executable so a weights file and its runtime become a single binary that runs on six operating systems with no installation.&lt;/p&gt;&#xA;&lt;p&gt;&lt;strong&gt;llamafile attacks distribution rather than serving: where Ollama, the family&amp;rsquo;s incumbent, gives your machine a model registry and a daemon, llamafile gives one person a file they can hand to another person, which is why it remains the only member here whose unit of delivery is an email attachment.&lt;/strong&gt;&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;What it is&#xA;    &lt;div id=&#34;what-it-is&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#what-it-is&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;A Mozilla Builders project launched November 2023, built by Justine Tunney (the Cosmopolitan Libc author) and now revamped by Mozilla.ai, licensed Apache-2.0 with its llama.cpp changes MIT so they can move upstream (26,190 stars, pushed 2026-10-07, as of 2026-10-07).&#xA;You point it at a GGUF weights file and get one cross-platform binary containing the weights, the inference engine, and a web UI, runnable on macOS, Linux, Windows, and the BSDs across six OS targets, on CPU or GPU, with no install step.&#xA;The same packaging ships whisperfile, a single-file speech-to-text and translation tool on whisper.cpp.&#xA;Release 0.10.6 (September 2026) moved to a new build system on the 0.10 line, and the project publishes pre-built llamafiles for popular models through Mozilla&amp;rsquo;s docs.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Status&#xA;    &lt;div id=&#34;status&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#status&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Active under Mozilla.ai stewardship: created 2023-09-10, 26,190 stars, release 0.10.6 published 2026-09-15, as of 2026-10-07.&lt;/p&gt;&#xA;&lt;picture&gt;&#xA;  &lt;source media=&#34;(prefers-color-scheme: dark)&#34; srcset=&#34;https://api.star-history.com/chart?repos=mozilla-ai/llamafile&amp;type=date&amp;theme=dark&amp;legend=top-left&#34; /&gt;&#xA;  &lt;source media=&#34;(prefers-color-scheme: light)&#34; srcset=&#34;https://api.star-history.com/chart?repos=mozilla-ai/llamafile&amp;type=date&amp;legend=top-left&#34; /&gt;&#xA;  &lt;img alt=&#34;Star History Chart&#34; src=&#34;https://api.star-history.com/chart?repos=mozilla-ai/llamafile&amp;type=date&amp;legend=top-left&#34; /&gt;&#xA;&lt;/picture&gt;&#xA;&lt;p&gt;The GitHub license field reads NOASSERTION because the repository combines Apache-2.0 with MIT-licensed llama.cpp changes; both Mozilla&amp;rsquo;s announcement and the README badge state Apache-2.0 as the project license.&#xA;Development cadence is slower than the local-inference family&amp;rsquo;s leaders (four releases in 2026), which is the trade of a format-stability project in a fast field.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Strengths&#xA;    &lt;div id=&#34;strengths&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#strengths&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;The lowest-friction distribution story in local inference: one file, no installer, no runtime setup, six OSes.&lt;/li&gt;&#xA;&lt;li&gt;Cosmopolitan&amp;rsquo;s reproduction promise, a given llamafile running the same weights the same way indefinitely, is unique in a category where toolchains churn monthly.&lt;/li&gt;&#xA;&lt;li&gt;GPU and dlopen support inside a single binary, so it is not a toy CPU-only path.&lt;/li&gt;&#xA;&lt;li&gt;Mozilla stewardship and Apache-2.0 licensing make it the least commercially exposed member of the family.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Cautions&#xA;    &lt;div id=&#34;cautions&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#cautions&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;Massive single binaries (weights plus engine) are clumsy to update: a new model version means a new multi-gigabyte file, where Ollama pulls a manifest.&lt;/li&gt;&#xA;&lt;li&gt;No model registry or discovery layer; you bring your own GGUF or use Mozilla&amp;rsquo;s pre-built set.&lt;/li&gt;&#xA;&lt;li&gt;Slower release cadence than llama.cpp itself, so fresh architecture support lags the upstream engine.&lt;/li&gt;&#xA;&lt;li&gt;The AVX2 requirement on x64 quick-install binaries excludes older Intel and AMD hardware and some VMs.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Pricing&#xA;    &lt;div id=&#34;pricing&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#pricing&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Free and open source (Apache-2.0, with llama.cpp changes MIT); there is no paid tier, so pricing does not apply.&#xA;Costs are your own hardware and the weights you bundle.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Compared to&#xA;    &lt;div id=&#34;compared-to&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#compared-to&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/model-access/ollama/&#34; &gt;Ollama&lt;/a&gt;: the registry-and-daemon incumbent; choose llamafile when the artifact must travel as one file, Ollama when a managed local model store matters more.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/model-access/magnitude/&#34; &gt;Magnitude&lt;/a&gt;: the self-optimizing inference engine; llamafile optimizes for distribution, Magnitude for kernel performance on your device.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/model-access/litellm/&#34; &gt;LiteLLM&lt;/a&gt;: the self-hosted gateway over cloud APIs; llamafile removes the cloud from the picture entirely.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Bottom line&#xA;    &lt;div id=&#34;bottom-line&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#bottom-line&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;&lt;strong&gt;Recommended for distributing a fixed model to non-technical recipients, offline environments, and archival use where the weights must stay runnable, and for the six-OS no-install demo.&lt;/strong&gt;&#xA;Not as a daily driver for people who churn models weekly, and not for serving agents that want a registry, quotas, or a daemon.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Changes&#xA;    &lt;div id=&#34;changes&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#changes&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;2026-10-07 - Created.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;See also&#xA;    &lt;div id=&#34;see-also&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#see-also&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/model-access/ollama/&#34; &gt;Ollama&lt;/a&gt; - the registry-and-daemon counterpart in the local-serving family&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/model-access/magnitude/&#34; &gt;Magnitude&lt;/a&gt; - the performance-focused local inference engine&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/model-access/litellm/&#34; &gt;LiteLLM&lt;/a&gt; - the gateway alternative for teams routing cloud providers instead&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/model-access/model-access-feature-matrix/&#34; &gt;Model Access Feature Matrix&lt;/a&gt; - the category comparison this note joins&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;References&#xA;    &lt;div id=&#34;references&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#references&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://github.com/mozilla-ai/llamafile&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=github.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://github.com/mozilla-ai/llamafile&lt;/a&gt; - repository, README, licensing badge, whisperfile, and the 0.10 build-system note (fetched 200, 2026-10-07)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://api.github.com/repos/mozilla-ai/llamafile&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=api.github.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://api.github.com/repos/mozilla-ai/llamafile&lt;/a&gt; - stars, created date, push date, and the NOASSERTION license field (fetched 200, 2026-10-07)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://hacks.mozilla.org/2023/11/introducing-llamafile/&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=hacks.mozilla.org&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://hacks.mozilla.org/2023/11/introducing-llamafile/&lt;/a&gt; - Mozilla&amp;rsquo;s launch announcement: the six-OS goal, Justine Tunney&amp;rsquo;s authorship, and the Apache-2.0 plus MIT split (fetched 200, 2026-10-07)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://docs.mozilla.ai/llamafile&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=docs.mozilla.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://docs.mozilla.ai/llamafile&lt;/a&gt; - the Mozilla.ai docs hub, pre-built llamafiles, and whisperfile surface (fetched 200, 2026-10-07)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://api.github.com/repos/mozilla-ai/llamafile/releases&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=api.github.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://api.github.com/repos/mozilla-ai/llamafile/releases&lt;/a&gt; - the 0.10.6 release of 2026-09-15 and the 2026 cadence (fetched 200, 2026-10-07)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://raw.githubusercontent.com/mozilla-ai/llamafile/HEAD/README.md&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=raw.githubusercontent.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://raw.githubusercontent.com/mozilla-ai/llamafile/HEAD/README.md&lt;/a&gt; - the llama.cpp plus Cosmopolitan architecture and the AVX2 x64 compatibility note (fetched 200, 2026-10-07)&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;</description>
      
    </item>
    
  </channel>
</rss>
