<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>benchmarking on tomrochette.com</title>
    <link>https://tomrochette.com/tags/benchmarking/</link>
    <description>Recent content in benchmarking on tomrochette.com</description>
    <generator>Hugo -- gohugo.io</generator>
    <language>en</language>
    <managingEditor>tom@tomrochette.com (Tom Rochette)</managingEditor>
    <webMaster>tom@tomrochette.com (Tom Rochette)</webMaster>
    <copyright>© 2026 Tom Rochette</copyright>
    <lastBuildDate>Wed, 23 Sep 2026 22:00:46 -0400</lastBuildDate><atom:link href="https://tomrochette.com/tags/benchmarking/index.xml" rel="self" type="application/rss+xml" />
    
    <item>
      <title>JevBench</title>
      <link>https://tomrochette.com/agents/hybrid-execution/jevbench/</link>
      <pubDate>Tue, 22 Sep 2026 00:00:00 +0000</pubDate>
      <author>tom@tomrochette.com (Tom Rochette)</author>
      <guid>https://tomrochette.com/agents/hybrid-execution/jevbench/</guid>
      <category>research-note</category><category>agent-curated</category><category>fully-ai-generated</category><category>llm=glm-5.3-flash</category><category>hybrid-execution</category><category>structured-outputs</category><category>system-one-models</category><category>decision-models</category><category>benchmarking</category><category>model-evaluation</category>
      <description>&lt;p&gt;JevBench is Benchmark Heaven&amp;rsquo;s MIT-licensed benchmark for Jev-class typed decision models: 534 English decisions per system, scored on chance-corrected intelligence, calibration, speed, and cost, with public items, frozen and hashed artifacts, and per-task outcomes checked into the repository.&#xA;Facts below verified as of 2026-09-22.&lt;/p&gt;&#xA;&lt;p&gt;&lt;strong&gt;This is the first third-party scoreboard for this category, and its headline reading, hosted Jev first at 74.4 with SemIf&amp;rsquo;s frozen-4B logit readout just 1.3 points behind on the same frozen items, is the closest thing the category has to independent verification, but it is one runner&amp;rsquo;s contested methodology, not a verdict.&lt;/strong&gt;&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;What it is&#xA;    &lt;div id=&#34;what-it-is&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#what-it-is&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;A benchmark harness by fstandhartinger under the Benchmark Heaven banner (benchmarkheaven.com), created 2026-09-19, MIT-licensed code, datasets, and scoring.&#xA;The v1.3.0 JevBench Score weights Intelligence, Calibration, Speed, and Cost at 25% each in a geometric mean, with a growing penalty below 50 intelligence so a cheap fast guesser cannot rank high.&#xA;Intelligence is measured above chance per tier (220 hard decisions written by Claude Opus 5 and GPT-5.6 Sol, cross-reviewed, frozen, and hashed before any system ran, plus easy, standard, and judge tiers, 534 decisions per system in total).&#xA;Calibration combines ECE on the hard tier with fidelity to exact gold distributions.&#xA;Cost is priced in dollars per 1,000 decisions, not per 1,000 tokens, which is the unit that actually matters for this contract.&#xA;The v1.2 board holds 48 ranked rows spanning hosted APIs, the open replicas, rerankers, and GLiNER-style classifiers, with classifier.dev listed as an honorable mention because its fast tier is Jev itself.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Status&#xA;    &lt;div id=&#34;status&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#status&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Active and four days old, as of 2026-09-22.&#xA;The repository was created 2026-09-19, pushed 2026-09-22, and shows 84 stars and 10 forks.&#xA;The Show HN thread (2026-09-22) sat at 92 points and 22 comments as of 2026-09-22, just under this category&amp;rsquo;s 100-point bar, and I state that gap explicitly.&#xA;What carries it over: it fills the exact gap every note in this category names (nobody independent had measured these systems against each other), and the artifacts are built to be re-run, not admired.&#xA;The author has 36 GitHub followers and no prior public footprint I could find, so I weight the methodology criticisms heavily.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Strengths&#xA;    &lt;div id=&#34;strengths&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#strengths&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;strong&gt;The scoreboard exists, and it is replayable: public items, frozen per-task outcomes, scoring code, and a script that rebuilds the final artifact from the frozen measurements.&lt;/strong&gt;&lt;/li&gt;&#xA;&lt;li&gt;The design anticipates this category&amp;rsquo;s specific tricks: chance-corrected intelligence defeats tiny-model majority-class games, calibration is scored against full distributions rather than argmax only, and cost is per decision.&lt;/li&gt;&#xA;&lt;li&gt;It measured the vendors too: Jev 1.13.0 (74.4), GPT-5.6 Luna (65.9 with the board&amp;rsquo;s best intelligence at 95.3), Gemini 3.1 Flash-Lite (60.1), and DeepSeek V4.1 Flash (57.5 with the best calibration at 96.7 and the worst cost at $0.5937 per 1,000 decisions).&lt;/li&gt;&#xA;&lt;li&gt;Combination experiments (confidence cascades, committees, best-of-n) are reported separately in RESULTS-COMBINATIONS.md, and none changed the ranked board, which is the negative result a lazy benchmark would have omitted.&lt;/li&gt;&#xA;&lt;li&gt;The limitations are stated in the launch post itself: English-only, latency from one German server, a disclosed times-two adjustment for self-hosted and demo endpoints, held-out prompts still reaching evaluated services, and roughly one-point gaps being noise.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Cautions&#xA;    &lt;div id=&#34;cautions&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#cautions&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;strong&gt;One runner built, ran, and scored everything, and the thread pushed back hard: one commenter said the results &amp;ldquo;do not seem to add up&amp;rdquo; and objected to the model set, another&amp;rsquo;s keysmash test of the slop-detector demo (86% confidence on keyboard noise) became the real finding, and a third read the site as vibecoded.&lt;/strong&gt;&lt;/li&gt;&#xA;&lt;li&gt;The times-two latency adjustment for self-hosted and demo endpoints is a disclosed assumption, not a measurement, and it moves every local row.&lt;/li&gt;&#xA;&lt;li&gt;The board mixes unlike things: djev is an inference method over DiffusionGemma rather than a trained model, several rows are author demo endpoints rather than production services, and some prices are announced preview prices.&lt;/li&gt;&#xA;&lt;li&gt;Small models are hypersensitive to option order on this suite too (one entrant scored 72% versus 21% on the same items with reversed options), so single-number rankings hide a real fragility.&lt;/li&gt;&#xA;&lt;li&gt;None of the category&amp;rsquo;s own notes can cite it as ground truth yet, because nobody has reproduced it.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Pricing&#xA;    &lt;div id=&#34;pricing&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#pricing&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Free and open: MIT-licensed harness, datasets, and scoring code, no hosted service and no paid tier.&#xA;The benchmark is free to run; the systems it scores bill at their own rates.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Compared to&#xA;    &lt;div id=&#34;compared-to&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#compared-to&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/hybrid-execution/nimble/&#34; &gt;Nimble&lt;/a&gt;: Bespoke&amp;rsquo;s PUBLIC_BENCHMARKS.md is the other independent yardstick, 13 human-labeled subsets measured against Jev only; JevBench covers the whole category but its gold labels are model-written, not human-labeled.&lt;/li&gt;&#xA;&lt;li&gt;The Jev workflow evals (evals.typesafe.ai): the vendor&amp;rsquo;s own scoreboard, which measures everything against an Astra-plus-Fable average and concedes the bias; JevBench is the outside answer to it.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/hybrid-execution/semif/&#34; &gt;SemIf&lt;/a&gt;: its RESULTS.md measures one frozen-model setup against TypeSafe&amp;rsquo;s published values; JevBench ran SemIf as an entrant and scored it second overall.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Bottom line&#xA;    &lt;div id=&#34;bottom-line&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#bottom-line&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;&lt;strong&gt;Recommended as the category&amp;rsquo;s first-stop scoreboard, read with the thread open in another tab, and as the artifact to re-run if you want to check any of its rows yourself.&lt;/strong&gt;&#xA;Not as grounds for a procurement decision, and not as a refutation of Jev: a 1.3-point lead over a frozen 4B model on one contested suite is a reason to demand better verification, not a conclusion.&#xA;The disagreeable claim I will defend: this benchmark&amp;rsquo;s most important number is not Jev&amp;rsquo;s 74.4, it is DeepSeek V4.1 Flash&amp;rsquo;s 96.7 calibration and Luna&amp;rsquo;s 95.3 intelligence, because they show the frontier models this category claims to beat are one score column away from competing, which should make everyone here uncomfortable.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Changes&#xA;    &lt;div id=&#34;changes&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#changes&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;2026-09-22 - Created from the entrant scan after the 2026-09-22 Show HN thread reached 92 points; accepted as category infrastructure despite the sub-100-point thread, with the author-standing gap and the thread&amp;rsquo;s methodology criticisms recorded.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;See also&#xA;    &lt;div id=&#34;see-also&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#see-also&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/hybrid-execution/jev/&#34; &gt;Jev&lt;/a&gt; - the closed model this benchmark ranks first, and the vendor whose antibenchmaxxing stance it answers&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/hybrid-execution/semif/&#34; &gt;SemIf&lt;/a&gt; - the frozen-model readout that lands 1.3 points behind Jev on this suite&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/hybrid-execution/nimble/&#34; &gt;Nimble&lt;/a&gt; - the human-labeled alternative yardstick covering Jev and one replica&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/hybrid-execution/hybrid-execution-feature-matrix/&#34; &gt;Hybrid Execution Feature Matrix&lt;/a&gt; - the category comparison this note now belongs to&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/model-selection-for-coding-tasks/&#34; &gt;Model Selection for Coding Tasks&lt;/a&gt; - where the decision layer you would benchmark this way gets chosen&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;References&#xA;    &lt;div id=&#34;references&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#references&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://github.com/fstandhartinger/jevbench&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=github.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://github.com/fstandhartinger/jevbench&lt;/a&gt; - repository: MIT, created 2026-09-19, 84 stars, 10 forks, pushed 2026-09-22 (GitHub API, as of 2026-09-22)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://raw.githubusercontent.com/fstandhartinger/jevbench/main/README.md&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=raw.githubusercontent.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://raw.githubusercontent.com/fstandhartinger/jevbench/main/README.md&lt;/a&gt; - the v1.3.0 score design, the 220 hard decisions frozen and hashed, the honorable-mention rule, and the option-order finding&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://raw.githubusercontent.com/fstandhartinger/jevbench/main/RESULTS-v1.2.md&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=raw.githubusercontent.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://raw.githubusercontent.com/fstandhartinger/jevbench/main/RESULTS-v1.2.md&lt;/a&gt; - the full 48-row board with per-axis scores, endpoints, and the times-two adjustment note&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://raw.githubusercontent.com/fstandhartinger/jevbench/main/results/v1.2/jevbench-v1.2-results.json&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=raw.githubusercontent.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://raw.githubusercontent.com/fstandhartinger/jevbench/main/results/v1.2/jevbench-v1.2-results.json&lt;/a&gt; - the frozen results artifact behind the board&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://news.ycombinator.com/item?id=49800574&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=news.ycombinator.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://news.ycombinator.com/item?id=49800574&lt;/a&gt; - the 92-point launch thread (2026-09-22, 22 comments), its methodology and model-set objections the critical source (fetched via the Algolia items API)&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;</description>
      
    </item>
    
  </channel>
</rss>
