<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>reasoning on tomrochette.com</title>
    <link>https://tomrochette.com/tags/reasoning/</link>
    <description>Recent content in reasoning on tomrochette.com</description>
    <generator>Hugo -- gohugo.io</generator>
    <language>en</language>
    <managingEditor>tom@tomrochette.com (Tom Rochette)</managingEditor>
    <webMaster>tom@tomrochette.com (Tom Rochette)</webMaster>
    <copyright>© 2026 Tom Rochette</copyright>
    <lastBuildDate>Tue, 06 Oct 2026 00:03:56 -0400</lastBuildDate><atom:link href="https://tomrochette.com/tags/reasoning/index.xml" rel="self" type="application/rss+xml" />
    
    <item>
      <title>Jeeves</title>
      <link>https://tomrochette.com/agents/hybrid-execution/jeeves/</link>
      <pubDate>Tue, 06 Oct 2026 00:00:00 +0000</pubDate>
      <author>tom@tomrochette.com (Tom Rochette)</author>
      <guid>https://tomrochette.com/agents/hybrid-execution/jeeves/</guid>
      <category>research-note</category><category>agent-curated</category><category>fully-ai-generated</category><category>llm=glm-5.3-flash</category><category>hybrid-execution</category><category>structured-outputs</category><category>system-one-models</category><category>decision-models</category><category>reasoning</category><category>open-weights</category>
      <description>&lt;p&gt;Jeeves is PostHog&amp;rsquo;s MIT-licensed research model, a 9B Jev-style decision model (a LoRA and pointer head on Qwen3.5-9B plus a small diffusion drafter) trained with SFT and CISPO to reason briefly before it answers typed questions.&lt;/p&gt;&#xA;&lt;p&gt;&lt;strong&gt;Jeeves is the first entrant to attack the wave&amp;rsquo;s core weakness, that single-pass decision models trade accuracy for speed, and its answer, think for about 3 seconds before deciding, beats hosted Jev&amp;rsquo;s published numbers on its own tests while conceding the speed the contract was built for.&lt;/strong&gt;&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;What it is&#xA;    &lt;div id=&#34;what-it-is&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#what-it-is&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;A research artifact from PostHog, the product-analytics company, acknowledging Kev as its inspiration: the same LoRA-plus-pointer-head architecture, plus a block-4 diffusion drafter, trained with supervised fine-tuning and CISPO so the model drafts a short reasoning chain and then reads out its answer.&#xA;It supports noul, choice, and score questions in one request through a Jev-compatible local server, runs on CUDA in bf16 or fp8 and on Apple Silicon through MPS, and ships full training code, train/dev/test data, and Apache-2.0 weights on Hugging Face.&#xA;Latency is the trade: about 0.3 s per request without thinking and a 3.3 s median with it on one H100 in fp8; on an M4 Mac the thinking runs at about 20 tokens per second.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Status&#xA;    &lt;div id=&#34;status&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#status&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;&lt;strong&gt;Days old and a research result, not a product.&lt;/strong&gt;&#xA;The repository was created 2026-09-29 and pushed 2026-10-01, with 411 stars and 21 forks as of 2026-10-06.&#xA;The launch thread (2026-09-29) reached 242 points, and Hugging Face shows 216 downloads and 5 likes.&#xA;Every benchmark in the README is self-run: on its held-out test split Jeeves scores 0.889 against Kev-9B&amp;rsquo;s published 0.822 and Jev&amp;rsquo;s published 0.857, and 0.935 against Jev&amp;rsquo;s 0.866 on JevBench&amp;rsquo;s 231 public items, while Jev keeps the transfer lead (0.800 against 0.746 on MMLU-Pro and buried state) and the same checkpoint without thinking drops to 0.804.&#xA;JevBench&amp;rsquo;s own board does not list Jeeves, so the public-tier numbers remain author-run.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Strengths&#xA;    &lt;div id=&#34;strengths&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#strengths&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;strong&gt;Reasoning composes with the decision contract: thinking lifts the same checkpoint from 0.804 to 0.889 on its test split, the clearest published answer yet to whether the single-pass accuracy ceiling is architectural or just untrained.&lt;/strong&gt;&lt;/li&gt;&#xA;&lt;li&gt;The unknown-question discipline beats the original: 0.055 of unknowable questions answered at p 0.9 or higher against Jev&amp;rsquo;s published 0.090, with better calibration error than Jev on public items (0.037 against 0.049).&lt;/li&gt;&#xA;&lt;li&gt;Everything needed to check it ships: training code, data, drafter, and a server, so the tables are re-runnable in a way most launch posts are not.&lt;/li&gt;&#xA;&lt;li&gt;Third-party adoption arrived within a week: Ollaya packages jeeves:9b in its library.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Cautions&#xA;    &lt;div id=&#34;cautions&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#cautions&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;strong&gt;Every comparison is self-run against published numbers measured on different samples, the caveat this wave keeps failing to escape, and the benchmark&amp;rsquo;s sealed tiers have never seen Jeeves.&lt;/strong&gt;&lt;/li&gt;&#xA;&lt;li&gt;Thinking costs roughly ten times the hosted contract&amp;rsquo;s latency (3.3 s median on an H100), which defeats the millisecond use case the category exists for; the result is a dial, not a free lunch.&lt;/li&gt;&#xA;&lt;li&gt;Apple silicon needs about 48 GB to avoid swapping with default caches.&lt;/li&gt;&#xA;&lt;li&gt;The thread&amp;rsquo;s skeptics landed the same two punches the whole wave absorbs: one called the primary use case &amp;ldquo;posting on social media about Jev&amp;rdquo;, and others asked why not just fine-tune a small model, with no independent evaluation yet on either side.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Pricing&#xA;    &lt;div id=&#34;pricing&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#pricing&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Free and open: MIT-licensed code and Apache-2.0 weights, no hosted service and no paid tier.&#xA;The cost is a CUDA GPU (or a large Mac) and the patience to wait out the thinking budget.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Compared to&#xA;    &lt;div id=&#34;compared-to&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#compared-to&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/hybrid-execution/kev/&#34; &gt;Kev&lt;/a&gt;: the acknowledged inspiration; Kev ships the locked-test eval discipline and the wider checkpoint family, Jeeves ships the reasoning result on top of the same architecture.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/hybrid-execution/jev/&#34; &gt;Jev&lt;/a&gt;: the closed original still wins transfer, latency, and independent scrutiny; Jeeves wins its own held-out test and the public JevBench items it ran.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/hybrid-execution/jeff/&#34; &gt;Jeff&lt;/a&gt;: the other second-generation student line; Jeff&amp;rsquo;s v1.3 adapters decide when a big model should think, Jeeves makes the small model think itself, and both keep the calibration-first framing.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Bottom line&#xA;    &lt;div id=&#34;bottom-line&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#bottom-line&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;&lt;strong&gt;Recommended for researchers of the decision contract who want to test whether brief reasoning generalizes, and for PostHog-scale teams with GPUs to spare.&lt;/strong&gt;&#xA;Not as a production dependency today: three weeks old, one company behind it, and no number a second party has confirmed.&#xA;The disagreeable claim I will defend: if brief reasoning generalizes across the wave, the single-pass architecture stops being the point and the contract becomes any calibrated classifier with a thinking budget, which is bad news for everyone selling speed as the moat.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Changes&#xA;    &lt;div id=&#34;changes&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#changes&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;2026-10-06 - Created from the entrant scan after the 2026-09-29 Show HN thread cleared the bar (242 points, PostHog standing, full training artifacts shipped).&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;See also&#xA;    &lt;div id=&#34;see-also&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#see-also&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/hybrid-execution/kev/&#34; &gt;Kev&lt;/a&gt; - the architecture Jeeves builds on and the eval discipline it should be judged against&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/hybrid-execution/jev/&#34; &gt;Jev&lt;/a&gt; - the closed model whose published numbers anchor every table here&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/hybrid-execution/jevbench/&#34; &gt;JevBench&lt;/a&gt; - the third-party scoreboard whose sealed tiers have not yet scored Jeeves&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/hybrid-execution/ollaya/&#34; &gt;Ollaya&lt;/a&gt; - the runtime that already packages jeeves:9b&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/hybrid-execution/hybrid-execution-feature-matrix/&#34; &gt;Hybrid Execution Feature Matrix&lt;/a&gt; - the category comparison this note joins&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;References&#xA;    &lt;div id=&#34;references&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#references&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://github.com/PostHog/jeeves&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=github.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://github.com/PostHog/jeeves&lt;/a&gt; - repository: MIT, created 2026-09-29, 411 stars, 21 forks, pushed 2026-10-01 (GitHub API, as of 2026-10-06)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://raw.githubusercontent.com/PostHog/jeeves/master/README.md&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=raw.githubusercontent.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://raw.githubusercontent.com/PostHog/jeeves/master/README.md&lt;/a&gt; - the benchmark tables, the LoRA-plus-pointer-head-plus-drafter architecture, CISPO training, latency table, and the Kev acknowledgement&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://news.ycombinator.com/item?id=49891290&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=news.ycombinator.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://news.ycombinator.com/item?id=49891290&lt;/a&gt; - the launch thread (242 points as of 2026-10-06, 2026-09-29), including the copycat and social-media skepticism&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://huggingface.co/PostHog/jeeves&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=huggingface.co&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://huggingface.co/PostHog/jeeves&lt;/a&gt; - weights: Apache-2.0, created 2026-09-29, 216 downloads and 5 likes as of 2026-10-06 (Hugging Face API)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://ollaya.dev/&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=ollaya.dev&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://ollaya.dev/&lt;/a&gt; - third-party adoption: jeeves:9b in the Ollaya library and its accuracy table&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;</description>
      
    </item>
    
  </channel>
</rss>
