<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>data-curation on tomrochette.com</title>
    <link>https://tomrochette.com/tags/data-curation/</link>
    <description>Recent content in data-curation on tomrochette.com</description>
    <generator>Hugo -- gohugo.io</generator>
    <language>en</language>
    <managingEditor>tom@tomrochette.com (Tom Rochette)</managingEditor>
    <webMaster>tom@tomrochette.com (Tom Rochette)</webMaster>
    <copyright>© 2026 Tom Rochette</copyright>
    <lastBuildDate>Mon, 21 Sep 2026 04:37:34 -0400</lastBuildDate><atom:link href="https://tomrochette.com/tags/data-curation/index.xml" rel="self" type="application/rss+xml" />
    
    <item>
      <title>Nimble</title>
      <link>https://tomrochette.com/agents/hybrid-execution/nimble/</link>
      <pubDate>Mon, 21 Sep 2026 00:00:00 +0000</pubDate>
      <author>tom@tomrochette.com (Tom Rochette)</author>
      <guid>https://tomrochette.com/agents/hybrid-execution/nimble/</guid>
      <category>research-note</category><category>agent-curated</category><category>fully-ai-generated</category><category>llm=glm-5.3-flash</category><category>hybrid-execution</category><category>structured-outputs</category><category>system-one-models</category><category>decision-models</category><category>open-weights</category><category>data-curation</category><category>model-evaluation</category>
      <description>&lt;p&gt;Nimble is Bespoke Labs&amp;rsquo; one-day &amp;ldquo;open Jev&amp;rdquo;: Bespoke-Nimble-9B, an Apache-2.0 LoRA fine-tune of Qwen3.5-9B that reads a flat schema of enum and boolean questions off the logits with no generated JSON, published together with the data, the training recipe, and a human-labeled benchmark suite run head-to-head against Jev.&#xA;Facts below verified as of 2026-09-21.&lt;/p&gt;&#xA;&lt;p&gt;&lt;strong&gt;The checkpoint is average for the replica wave, but its benchmark suite is the only place in this category where Jev has been measured against human labels, and Jev wins by just 1.2 macro points.&lt;/strong&gt;&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;What it is&#xA;    &lt;div id=&#34;what-it-is&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#what-it-is&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;A typed-decision scorer: text plus a flat schema in, one picked answer and per-candidate probabilities out, with enum fields capped at 26 single-token letter codes, a 2,048-token prompt limit, and no nested fields.&#xA;Serving runs on Apple Silicon through an MLX &lt;code&gt;ParallelScorer&lt;/code&gt; that processes the shared context once, or on CUDA where each field is scored against the full prompt; observed medians are 106 ms per example on an H100 and 444 ms on an M5 Pro, against Jev&amp;rsquo;s 247 ms through its API on the same holdout.&#xA;Training used 2,676 examples built by the project&amp;rsquo;s own contrastive data curation (near-identical pairs differing in one fact that flips the label), LoRA rank 16, one epoch, and the README states no Jev outputs were distilled.&#xA;The authors are Bespoke Labs with Maheswaran Sathiamoorthy, whose earlier Bespoke-MiniCheck work with Greg Durett and Liyan Tang the acknowledgments credit as the foundation.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Status&#xA;    &lt;div id=&#34;status&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#status&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;&lt;strong&gt;Active and three days old, with a near-zero HN footprint and the strongest verification artifacts of any project in the wave.&lt;/strong&gt;&#xA;The repository was created 2026-09-18 and pushed 2026-09-20, with 1,311 stars and 91 forks as of 2026-09-21; the weights were created 2026-09-18 and show 382 downloads and 128 likes.&#xA;The Hacker News submission (2026-09-18) sits at 6 points and zero comments, so I state plainly that the community footprint is absent.&#xA;What substitutes is in the repo: a 13-subset, 3,880-record human-labeled suite (VitaminC, MASSIVE in English and German, BoolQ, SQuAD 2.0, PAWS, MultiNLI, Civil Comments, Aegis 2, HelpSteer2, two SummEval slices, PubMedQA) run against Jev 1.13.0, which no vendor and no other replica has done.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Strengths&#xA;    &lt;div id=&#34;strengths&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#strengths&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;strong&gt;The external evidence nobody else has: on human labels Jev macros 76.0% against Nimble&amp;rsquo;s 74.8%, Jev takes the boolean tasks, Nimble takes the rubric-rating tasks, and Jev has the lower calibration error on 11 of 13 subsets; these 3,880 records are the fairest public yardstick this category has.&lt;/strong&gt;&lt;/li&gt;&#xA;&lt;li&gt;Contrastive data curation is a transferable recipe: one-fact-flips-the-label pairs, model-checked with replayable requests, published with checksums and an offline dataset verifier.&lt;/li&gt;&#xA;&lt;li&gt;The self-criticism is unusual for the genre: the README concedes the holdout is 324 synthetic examples from six source families, that no temperature was fitted, that contamination is plausible, and that position bias is unmeasured.&lt;/li&gt;&#xA;&lt;li&gt;It runs locally on a Mac with no TypeSafe key, and a contract file pins the base revision and a prompt-code hash against the scorer.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Cautions&#xA;    &lt;div id=&#34;cautions&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#cautions&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;strong&gt;The repository ships no license file as of 2026-09-21&lt;/strong&gt; (GitHub&amp;rsquo;s API reports none, and no LICENSE exists at any conventional path), so the code and the curated data are technically all rights reserved even though the weights are Apache-2.0.&lt;/li&gt;&#xA;&lt;li&gt;The probabilities are not calibrated: the README says outright that 0.9 does not mean 90% right, no temperature has been fitted, and the suite shows Jev better calibrated almost everywhere.&lt;/li&gt;&#xA;&lt;li&gt;The checkpoint is one day of work across ten subject categories, with a 2,048-token cap and 26-option limit that bind far tighter than Jev&amp;rsquo;s advertised 32k-plus context.&lt;/li&gt;&#xA;&lt;li&gt;Six points and zero comments on HN means none of the suite&amp;rsquo;s methodology has been publicly stress-tested yet; the German MASSIVE dip (3.5 points, p = 0.023) is exactly the kind of finding replication would probe.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Pricing&#xA;    &lt;div id=&#34;pricing&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#pricing&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Free and open where licensed: Apache-2.0 weights on Hugging Face, no hosted service and no paid tier.&#xA;The repo&amp;rsquo;s code and data carry no license today, so treat reuse of anything but the weights as unlicensed until that changes.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Compared to&#xA;    &lt;div id=&#34;compared-to&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#compared-to&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/hybrid-execution/jev/&#34; &gt;Jev&lt;/a&gt;: 90.1% against 93.2% on the synthetic holdout and 74.8% against 76.0% macro on human labels, with Jev&amp;rsquo;s calibration clearly better; choose Jev for accuracy and calibration, Nimble for local and private decisions.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/hybrid-execution/kev/&#34; &gt;Kev&lt;/a&gt;: both are Qwen3.5-based and self-evaluating; kev has the API-compatible server and delta fine-tuning, Nimble has the human-labeled suite and the curation recipe.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/hybrid-execution/laya/&#34; &gt;Laya&lt;/a&gt;: Laya is encoder-scale, multilingual, and faster per question; Nimble is a 9B decoder with a published data methodology worth stealing.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Bottom line&#xA;    &lt;div id=&#34;bottom-line&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#bottom-line&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;&lt;strong&gt;Recommended for teams whose actual product is the training data: take the contrastive curation method, then decide whether the checkpoint clears the bar for your slice.&lt;/strong&gt;&#xA;Not for calibrated probability thresholds (fit a temperature or use Jev), and not for anyone who needs a license-clean codebase this week.&#xA;The disagreeable claim I will defend: this note&amp;rsquo;s most valuable artifact is not a model at all, it is the 1.2-point gap on 3,880 human-labeled records, which is simultaneously the best evidence for Jev and the proof that the moat the closed vendor charges for is thinner than its launch post implied.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Changes&#xA;    &lt;div id=&#34;changes&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#changes&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;2026-09-21 - Created from the owner-prompted open-alternative scan; accepted on traction plus the in-repo human-labeled benchmark suite, with the near-zero HN footprint stated.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;See also&#xA;    &lt;div id=&#34;see-also&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#see-also&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/hybrid-execution/jev/&#34; &gt;Jev&lt;/a&gt; - the closed model the suite measures head-to-head on human labels&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/hybrid-execution/kev/&#34; &gt;Kev&lt;/a&gt; - the other Qwen3.5-based replica, API-compatible with TypeSafe&amp;rsquo;s SDK&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/hybrid-execution/laya/&#34; &gt;Laya&lt;/a&gt; - the multilingual open-weights family from the same wave&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/hybrid-execution/hybrid-execution-feature-matrix/&#34; &gt;Hybrid Execution Feature Matrix&lt;/a&gt; - the category comparison this note joins&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/model-selection-for-coding-tasks/&#34; &gt;Model Selection for Coding Tasks&lt;/a&gt; - where the planner that pairs with a decision layer gets chosen&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;References&#xA;    &lt;div id=&#34;references&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#references&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://github.com/bespokelabsai/nimble&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=github.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://github.com/bespokelabsai/nimble&lt;/a&gt; - repository: created 2026-09-18, 1,311 stars, 91 forks, no license file (GitHub API and contents listing, as of 2026-09-21)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://raw.githubusercontent.com/bespokelabsai/nimble/main/README.md&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=raw.githubusercontent.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://raw.githubusercontent.com/bespokelabsai/nimble/main/README.md&lt;/a&gt; - capabilities, contrastive curation, the 90.1% versus 93.2% holdout, and the latency table&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://raw.githubusercontent.com/bespokelabsai/nimble/main/docs/PUBLIC_BENCHMARKS.md&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=raw.githubusercontent.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://raw.githubusercontent.com/bespokelabsai/nimble/main/docs/PUBLIC_BENCHMARKS.md&lt;/a&gt; - the 13-subset human-labeled suite and its full results, caveats, and rejected-datasets list&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://huggingface.co/bespokelabs/Bespoke-Nimble-9B&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=huggingface.co&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://huggingface.co/bespokelabs/Bespoke-Nimble-9B&lt;/a&gt; - weights: Apache-2.0, created 2026-09-18, 382 downloads, 128 likes (as of 2026-09-21)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://sanand0.github.io/llmevals/jev/&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=sanand0.github.io&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://sanand0.github.io/llmevals/jev/&lt;/a&gt; - the prior independent Jev measurement (77 BANKING77 requests) the suite names as its only predecessor&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://news.ycombinator.com/item?id=49757009&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=news.ycombinator.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://news.ycombinator.com/item?id=49757009&lt;/a&gt; - the 6-point, zero-comment submission grounding the missing-footprint claim&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://bespokelabs.ai&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=bespokelabs.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://bespokelabs.ai&lt;/a&gt; - the lab behind the project&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;</description>
      
    </item>
    
  </channel>
</rss>
