
Jeeves is PostHog's MIT-licensed research model, a 9B Jev-style decision model (a LoRA and pointer head on Qwen3.5-9B plus a small diffusion drafter) trained with SFT and CISPO to reason briefly before it answers typed questions.

**Jeeves is the first entrant to attack the wave's core weakness, that single-pass decision models trade accuracy for speed, and its answer, think for about 3 seconds before deciding, beats hosted Jev's published numbers on its own tests while conceding the speed the contract was built for.**

## What it is

A research artifact from PostHog, the product-analytics company, acknowledging Kev as its inspiration: the same LoRA-plus-pointer-head architecture, plus a block-4 diffusion drafter, trained with supervised fine-tuning and CISPO so the model drafts a short reasoning chain and then reads out its answer.
It supports noul, choice, and score questions in one request through a Jev-compatible local server, runs on CUDA in bf16 or fp8 and on Apple Silicon through MPS, and ships full training code, train/dev/test data, and Apache-2.0 weights on Hugging Face.
Latency is the trade: about 0.3 s per request without thinking and a 3.3 s median with it on one H100 in fp8; on an M4 Mac the thinking runs at about 20 tokens per second.

## Status

**Days old and a research result, not a product.**
The repository was created 2026-09-29 and pushed 2026-10-01, with 411 stars and 21 forks as of 2026-10-06.
The launch thread (2026-09-29) reached 242 points, and Hugging Face shows 216 downloads and 5 likes.
Every benchmark in the README is self-run: on its held-out test split Jeeves scores 0.889 against Kev-9B's published 0.822 and Jev's published 0.857, and 0.935 against Jev's 0.866 on JevBench's 231 public items, while Jev keeps the transfer lead (0.800 against 0.746 on MMLU-Pro and buried state) and the same checkpoint without thinking drops to 0.804.
JevBench's own board does not list Jeeves, so the public-tier numbers remain author-run.

## Strengths

- **Reasoning composes with the decision contract: thinking lifts the same checkpoint from 0.804 to 0.889 on its test split, the clearest published answer yet to whether the single-pass accuracy ceiling is architectural or just untrained.**
- The unknown-question discipline beats the original: 0.055 of unknowable questions answered at p 0.9 or higher against Jev's published 0.090, with better calibration error than Jev on public items (0.037 against 0.049).
- Everything needed to check it ships: training code, data, drafter, and a server, so the tables are re-runnable in a way most launch posts are not.
- Third-party adoption arrived within a week: Ollaya packages jeeves:9b in its library.

## Cautions

- **Every comparison is self-run against published numbers measured on different samples, the caveat this wave keeps failing to escape, and the benchmark's sealed tiers have never seen Jeeves.**
- Thinking costs roughly ten times the hosted contract's latency (3.3 s median on an H100), which defeats the millisecond use case the category exists for; the result is a dial, not a free lunch.
- Apple silicon needs about 48 GB to avoid swapping with default caches.
- The thread's skeptics landed the same two punches the whole wave absorbs: one called the primary use case "posting on social media about Jev", and others asked why not just fine-tune a small model, with no independent evaluation yet on either side.

## Pricing

Free and open: MIT-licensed code and Apache-2.0 weights, no hosted service and no paid tier.
The cost is a CUDA GPU (or a large Mac) and the patience to wait out the thinking budget.

## Compared to

- [Kev](../kev/index.md): the acknowledged inspiration; Kev ships the locked-test eval discipline and the wider checkpoint family, Jeeves ships the reasoning result on top of the same architecture.
- [Jev](../jev/index.md): the closed original still wins transfer, latency, and independent scrutiny; Jeeves wins its own held-out test and the public JevBench items it ran.
- [Jeff](../jeff/index.md): the other second-generation student line; Jeff's v1.3 adapters decide when a big model should think, Jeeves makes the small model think itself, and both keep the calibration-first framing.

## Bottom line

**Recommended for researchers of the decision contract who want to test whether brief reasoning generalizes, and for PostHog-scale teams with GPUs to spare.**
Not as a production dependency today: three weeks old, one company behind it, and no number a second party has confirmed.
The disagreeable claim I will defend: if brief reasoning generalizes across the wave, the single-pass architecture stops being the point and the contract becomes any calibrated classifier with a thinking budget, which is bad news for everyone selling speed as the moat.

## Changes

- 2026-10-06 - Created from the entrant scan after the 2026-09-29 Show HN thread cleared the bar (242 points, PostHog standing, full training artifacts shipped).

## See also

- [Kev](../kev/index.md) - the architecture Jeeves builds on and the eval discipline it should be judged against
- [Jev](../jev/index.md) - the closed model whose published numbers anchor every table here
- [JevBench](../jevbench/index.md) - the third-party scoreboard whose sealed tiers have not yet scored Jeeves
- [Ollaya](../ollaya/index.md) - the runtime that already packages jeeves:9b
- [Hybrid Execution Feature Matrix](../hybrid-execution-feature-matrix/index.md) - the category comparison this note joins

## References

- https://github.com/PostHog/jeeves - repository: MIT, created 2026-09-29, 411 stars, 21 forks, pushed 2026-10-01 (GitHub API, as of 2026-10-06)
- https://raw.githubusercontent.com/PostHog/jeeves/master/README.md - the benchmark tables, the LoRA-plus-pointer-head-plus-drafter architecture, CISPO training, latency table, and the Kev acknowledgement
- https://news.ycombinator.com/item?id=49891290 - the launch thread (242 points as of 2026-10-06, 2026-09-29), including the copycat and social-media skepticism
- https://huggingface.co/PostHog/jeeves - weights: Apache-2.0, created 2026-09-29, 216 downloads and 5 likes as of 2026-10-06 (Hugging Face API)
- https://ollaya.dev/ - third-party adoption: jeeves:9b in the Ollaya library and its accuracy table
