Pavan Kumar T V

CTO | Technology Leader

By ·

Whose GPU?

Series: Part 1 is what a decision model like Jev is good for, Part 2 what it does to a GPU, Part 3 why deciding is not predicting, and Part 4 builds a predictor and grades it. This is Part 5, on the new wave of open decision models, the hardware behind their numbers, and how to work out what fits on yours.

Two new decision models landed this week.

Cloudflare shipped Clef and Clef-flash on 1 October. Strands Labs has Strands Decider, a 2B model with the whole research log in the repo.

The numbers on the posters: 38.8 ms. 115 ms. 209 ms.

My 4B model takes most of a second per email on a T4.

My first thought was that I'm doing it wrong. My second thought was better.

Whose GPU was that?

So I downloaded them, opened the weight files, and ran six on my laptop. This post is what I found, and the arithmetic you can use to check any model before you download it.

Words You'll Need

Part 2 covers tokens, prefill, quantization and the T4. Part 3 covers hidden states. These are the new ones.

  • Parameter (weight). One number inside the model. "2B" means roughly two billion of them.
  • Forward pass. One trip through the model. Text goes in, every layer runs once, numbers come out.
  • Embedding table. A lookup table with one row of numbers per token in the vocabulary. The model's front door.
  • Backbone (torso). The stack of layers in the middle. Where the actual computing happens.
  • Head. The small piece on the end that turns the backbone's output into an answer.
  • Base model. The general model a decision model was built on. Somebody else trained it; the decision model adds a little on top.
  • Adapter (LoRA). A small set of extra weights trained on top of a frozen base model. Cheaper than retraining the whole thing.
  • VRAM. The GPU's own memory. The weights have to fit in it before anything else happens.
  • bf16, fp16, fp32. Number formats. 16 bits is two bytes per weight, 32 bits is four.
  • GGUF. The file format llama.cpp reads. Usually holds a quantized model.
  • Median, p90, p95. The median is the middle request. The p95 is the slow one, the request 19 out of 20 are faster than. Users remember the slow one.
  • Cold and warm. Cold is the first time the machine sees a request of that shape. Warm is every time after.

Two New Models

Same three questions as Jev. Pick one, put it on a scale, yes or no. Both are Apache 2.0 with open weights. Both speak the same request format.

                   Clef          Clef-flash     Strands Decider
  ────────────────────────────────────────────────────────────
  Base model       Qwen3.8-27B   Qwen3.5-9B     Qwen3.5-2B-Base
  Parameters       27.4B         9.4B           1.9B
  Head             joint head    joint head     pointer head
  Median latency   209.3 ms      38.8 ms        115 ms

Two things in there made me smile.

In Part 3 I said the letter softmax is the worst place to read a decision from. Neither of these reads letters. Both put their own head on the backbone's hidden states.

In Part 4 I said the model's probabilities needed grading. Clef trains with a Brier loss. Strands reports calibration error next to accuracy.

Good. The field moved the right way.

Now the latency row. To read it, you need to know what one decision costs.

One Forward Pass

A chat model runs once per word it writes. A decision model runs once, full stop. That's the whole trick, and it's why the numbers can be so small.

Here's what that one pass looks like inside Strands Decider. Every number below is from the config and weight files I downloaded.

  your text + the question
            │
            ▼
  tokenizer ──▶ token ids
            │
            ▼
  embedding table         248,320 rows x 2,048 numbers
  (a lookup, no maths)    0.51B weights
            │
            ▼
  24 layers               3 linear-attention layers, then
  (the backbone)          1 full-attention layer, six times
                          1.37B weights
            │             every token passes through every one
            ▼
  final hidden states     2,048 numbers per token
            │
            ▼
  pointer head            1,053,184 weights
            │
            ▼
  one probability per option

Two different costs hide in that picture.

Memory is paid by every weight, used or not. The embedding table is half a billion numbers and it has to sit somewhere.

Time is paid by the backbone. The table is a lookup: one row per token, no multiplying. The layers are where every token gets multiplied through 1.37 billion weights.

So time grows with two things. How many backbone weights. How many tokens you sent. Double the text, roughly double the wait. Nothing in the output is long, so nothing else matters.

It also means the usual GPU tricks don't transfer. Batching helps a chat model because writing one word at a time leaves the card waiting on memory. A decision model never writes. It's all reading, and reading already keeps the card busy. I tried batching in Part 2 and got nothing for it.

Clef-flash is the same shape, bigger. 32 layers instead of 24. 4,096 numbers per token instead of 2,048. And a much heavier head: 122 million weights, a small 4-layer transformer of its own, against Strands' one million.

What "2B" Actually Means

Here's the thing: the number in the model's name is a label, not a count.

I opened the Qwen3.5-2B-Base weight file and added it up.

  Qwen3.5-2B-Base, as downloaded        Weights
  ─────────────────────────────────────────────
  Backbone layers                        1.37B
  Token embedding table                  0.51B
  Vision tower (for images)              0.33B
  Extra prediction layers                0.06B
  ─────────────────────────────────────────────
  Total in the file                      2.27B
  Text only                              1.88B

Three things to take from that.

  • The file is bigger than the name. 2.27 billion weights, 4.55 GB. Strands calls its model 1.9B because it only uses the text part.
  • A quarter of a "2B" text model is a lookup table. 0.51B of 1.88B. It costs memory and almost no time. The part that does the work is 1.37B.
  • Small models carry the same vocabulary as big ones. All of these use 248,320 tokens. The table shrinks only as fast as the row width, so the smaller the model, the bigger the table's share.

The 9B has a twist. In the 2B, one table does two jobs: turn tokens into numbers on the way in, and numbers back into tokens on the way out. In Clef-flash they're two separate tables, about a billion weights each.

That's 2 of its 9.4 billion. And the output one is the language head, the thing a decision model was built to skip.

One more label worth reading: Base. Strands builds on the Base model, the raw one that was never taught to chat. For a model that will never write a sentence, that's the right starting point.

How to Size a Model

You can do this before downloading anything. One multiplication.

Size in GB = weights in billions x bits per weight / 8

Check it against the real files:

  • Clef-flash: 9.41B x 16 bits / 8 = 18.8 GB. Add the 0.24 GB head. The repo is 19.08 GB.
  • Clef: 27.36B x 16 / 8 = 54.7 GB. The repo is 54.99 GB.
  • Strands, text only: 1.88B x 16 / 8 = 3.8 GB.

Now the part people get wrong. "4-bit" is not 4 bits.

A team at Livesport published a llama.cpp build of Clef-flash with the sizes of every variant they tried. Divide each by the 16-bit size and you get the real bits per weight.

  Format            Clef-flash on the GPU   Real bits per weight
  ──────────────────────────────────────────────────────────────
  bf16                    15.9 GB                 16
  Q8_0                     8.44 GB           about 8.5
  Q5_K_S                   5.79 GB           about 5.8
  Q4_K_M                   5.05 to 5.26 GB   about 5.1 to 5.3

A "4-bit" model costs a bit over 5. Some layers are kept at higher precision on purpose, because squeezing them hurts too much.

Then add what the weights don't cover:

  • The head. 0.24 GB for Clef-flash. Nothing for Strands.
  • Working memory for the text. It grows with prompt length. Livesport measured 7.33 GB used on an 8 GB card after maximum-length requests, with 5.57 GB of that being weights.
  • The scratch buffer. llama.cpp reads text in chunks, and the buffer scales with the chunk. On my 4B: 495 MiB at 512 tokens, 1,980 MiB at 2,048. At 8,192 it asked for 7,920 MiB and the process died.
  • Per-request state. In Part 4 one wrong setting reserved 12.8 GB of it and killed my process.

Rule of thumb from those numbers: leave a quarter of the card free.

If you want the exact count, the weight file tells you. A safetensors file starts with a plain list of every tensor and its shape:

import json, struct

f = open("model.safetensors", "rb")
n = struct.unpack("<Q", f.read(8))[0]
header = json.loads(f.read(n))
header.pop("__metadata__", None)

total = 0
for name, t in header.items():
    count = 1
    for dim in t["shape"]:
        count *= dim
    total += count
print(total)

That's how I got every count in this post. It reads a few kilobytes, not the whole file.

The Cards

  Card               Memory    Generation      Runs bf16 natively
  ───────────────────────────────────────────────────────────────
  Quadro RTX 4000      8 GB    Turing, 2018    no
  T4                  16 GB    Turing, 2018    no
  RTX 3090            24 GB    Ampere, 2020    yes
  A100             40/80 GB    Ampere, 2020    yes
  H100                80 GB    Hopper, 2022    yes
  H200               141 GB    Hopper, 2023    yes

The T4 is the card every cloud rents cheaply. The 3090 is a gaming card that lives in home labs. The bottom three are rented by the hour and nobody puts one under a desk.

Put the sizing rule and the card list together:

  Weights as published (bf16)     T4      3090    H100    H200
                                  16 GB   24 GB   80 GB   141 GB
  ─────────────────────────────────────────────────────────────
  Strands Decider    3.8 GB       fits    fits    fits    fits
  Clef-flash        19.1 GB       no      fits    fits    fits
  Clef              55.0 GB       no      no      fits    fits

Clef-flash, the small fast one, is 19.1 GB. A T4 has 16.

Clef is 55 GB. Even squeezed to a "4-bit" build, the sum comes to 17 or 18 GB. Still over a T4.

Open weights are not the same as local.

The Number on the Poster

The Cloudflare post has a latency table. Clef at 209.3 ms median, Clef-flash at 38.8 ms, Jev at 524.1 ms, "across the 43 eval benchmarks that we ran".

It doesn't say what card. It doesn't say how long the prompts were.

So I opened the model card. One line under Usage:

That's the only hardware the release names. To be fair, the card doesn't say the latency rows came from that H200. It doesn't say they came from anywhere. But nothing in the release mentions a smaller card.

You can test the claim with the forward pass. Same back-of-envelope as Part 2: about 2 FLOPs per weight per token, and a T4 that peaks around 65 TFLOPS.

How many tokens could a T4 read inside each published median, running at a peak it never reaches?

                Published     Tokens a T4 could read
                median        in that time, at peak
  ──────────────────────────────────────────────────
  Clef-flash     38.8 ms      about 130
  Clef          209.3 ms      about 250

My emails run 80 to 1,100 tokens.

I used every weight in that sum. Leave out the lookup tables and the numbers move by a quarter or so. The answer doesn't.

38.8 ms is not a T4 number. It can't be.

What a Small Card Actually Does

I don't have to guess any more, because somebody measured it.

Livesport run Clef-flash in production on a Quadro RTX 4000. Same chip generation as the T4, half the memory. One request at a time:

  Text length          1 question   5 questions
  ─────────────────────────────────────────────
  about 600 tokens       0.57 s       0.86 s
  about 2,000 tokens     1.75 s       2.02 s
  about 7,400 tokens     6.58 s       6.97 s
  16,384 tokens         15.2 s       15.9 s

About 1,100 tokens a second. The backbone is 98% of the time.

Read that table against the forward pass. Time climbs with tokens, almost in a straight line. Four extra questions cost 0.3 seconds, because they ride on the same pass.

So on a small card, a 600-token decision is 570 ms. The poster says 38.8.

Nobody lied. They're different machines.

On My Laptop

Then I ran them. An M3 Pro with 18 GB, nothing else special.

Strands Decider as published, on the Mac's GPU. Clef-flash as the Livesport GGUF through llama.cpp, with Cloudflare's head on top. Then four more from the same wave, all on the Mac's GPU: K2-Type, RSI-Jev, Sol from Decision 2.0 and Gero-4B. One model at a time. The laptop runs hot enough already.

Two kinds of text: my work emails, and the 779 public GitHub issues from Part 4.

                        Emails              GitHub issues
                        median    p90       median    p90
  ──────────────────────────────────────────────────────────
  K2-Type 0.9B          100 ms    121 ms    0.6 s     0.9 s
  Strands Decider 2B    253 ms    290 ms    1.2 s     2.0 s
  RSI-Jev 2B            321 ms    394 ms    1.8 s     3.1 s
  Sol 2B                458 ms    482 ms    2.0 s     3.2 s
  Gero-4B               1.3 s     1.7 s     3.6 s     5.9 s
  Clef-flash 9B         1.4 s     1.6 s     4.7 s     6.9 s

Every issue column is all 779, except Gero's, which is the first 199. I stopped that run early, more on that below.

The Clef-flash run finished. Its slowest issue took 10.6 seconds. Partway through, the laptop got hot enough that I paused it, shut down everything else on the machine, and started again.

That's the honest local result for the 9B. It loads. It answers. It's nearly five seconds a decision, and you'll hear the fans.

Gero is the odd one. It isn't the biggest, but it scores every option as its own pass through the model. A yes-or-no question costs two forward passes. A pick-one-of-three costs three.

One caution on that table. It compares stacks, not just models. Clef-flash runs in llama.cpp with a Python head bolted on. The rest run in PyTorch. K2-Type, RSI-Jev and Sol ran in 32-bit. Gero ran in 16-bit, because a 4B in 32-bit is 16 GB and the laptop has 18. Some of the gaps are model size. Some of it is plumbing.

The 2B is a different story. A quarter of a second on short emails is usable. But the same model, same laptop, took four to five times longer on the issues. The only thing that changed was the text.

Strands' own docs say the same. Their headline is 115 ms on an RTX 3090, as a median over a benchmark where 136 of 231 tasks are under 300 tokens. Then they print this:

  Prompt tokens    3090      M3 Pro, warm   M3 Pro, first time
  ────────────────────────────────────────────────────────────
  under 300        112 ms      153 ms          310 ms
  300 to 1,000     115 ms      449 ms        1,644 ms
  1,000 to 2,500   228 ms    2,172 ms        2,851 ms
  2,500 to 5,000   286 ms    2,514 ms        3,836 ms
  ────────────────────────────────────────────────────────────
  all, median      115 ms      234 ms          622 ms
  all, p95         299 ms    2,628 ms        3,862 ms

My numbers sit inside theirs. I have nothing but respect for that repo. They even write that the Mac figures are "not a latency claim for other hardware". But nobody quotes the footnote. They quote 115 ms.

Why the CPU Matters

GPU talk skips the machine most teams actually have.

  • It's already there. Every server has a CPU. No quota request, no budget meeting.
  • RAM is the cheap memory. 64 GB of RAM is an ordinary server. 64 GB of VRAM is an H100.
  • The data stays put. Some text can't leave the building. A CPU in the building is the simplest answer to that.
  • Part of the model lives there anyway. In the Livesport build, the 1.08 GB embedding table and the 2 GB language head stay in ordinary RAM. In both models I ran, the head can run on the CPU. Only the backbone needs the card.

What it costs is time. Same forward pass, slower arithmetic.

From Strands' docs, the 2B on an M3 Pro CPU:

  • 256-token text, one question: 3.3 seconds
  • 1,024-token text, one question: 5.4 seconds
  • Peak memory: 8.3 GB against 5.4 GB on the GPU, because on CPU it runs in 32-bit and every weight doubles

And mine. My 4B at 4 bits, a short text with three questions, same laptop:

  M3 Pro              Per request
  ─────────────────────────────────
  GPU (Metal)         about 0.76 s
  CPU, warm           3.0 to 4.2 s
  CPU, first run      7.7 s

Four to five times slower. Not forty.

Do the sum before you dismiss it. Three seconds a decision is 1,200 an hour. Almost 29,000 a day, on one box, for free.

That's not a text box. It's a perfectly good nightly batch.

Can You Squeeze It?

In my first draft of this post I wrote that nobody had published a quantized Clef. I was wrong. There's the Livesport build, bartowski's, an official ggml-org one, and an open pull request to make llama.cpp run Clef natively.

I also assumed squeezing would wreck the probabilities. Strands' hardware notes warn about exactly that:

Livesport measured it. 900 questions plus 24 long documents, every variant compared with the full 16-bit model.

  Variant                        On GPU    Same answer   Probability
                                           as bf16       shift
  ──────────────────────────────────────────────────────────────────
  Q8_0, their recipe             7.49 GB     99.9%        0.0018
  Q5_K_M, their recipe           5.57 GB     99.8%        0.0036
  Q4_K_M, careful public build   5.26 GB     99.2%        0.0065
  Q4_K_M, plain                  5.05 GB     98.6%        0.0126

Both halves of the story are in that table.

Done carefully, a 5.57 GB build gives the same answer 998 times in 1,000 and moves the probability by a third of a percentage point. Done carelessly, the shift is three and a half times bigger for half a gigabyte saved.

Their trick is worth knowing. The language head is useless to a decision model, so they crushed it to 2 bits and spent the saved memory on the backbone.

That only works if you know which weights your forward pass touches.

The warning still stands. It just has an answer now: measure it, on your own data.

Fast Is Not Right

Fitting is half of it. So I graded them on the Part 4 question. Will this GitHub issue be closed within seven days? Same 779 issues, no training, just ask.

  Zero-shot, all 779 issues     AUC   Brier     ECE
  ─────────────────────────────────────────────────
  Base rate                   0.500   0.195   0.108
  My 4B, ask (Part 4)         0.613   0.190   0.107
  Clef-flash 9B, 5-bit        0.608   0.194   0.089
  RSI-Jev 2B                  0.600   0.180   0.026
  Sol 2B                      0.569   0.185   0.058
  K2-Type 0.9B                0.547   0.246   0.237
  Strands Decider 2B          0.500   0.186   0.044

Start with Strands. Its probabilities are well behaved. It said about 27% on average and the real rate was 24%.

And its ranking is a coin flip. AUC 0.500. It said "yes" outright once in 779 issues.

A model that says "about a quarter" to everything is calibrated and useless. That's the trap in a calibration score: sitting on the base rate earns a good one. Always read it next to a ranking number.

I haven't worked out why. It could be the wording of my question. It could be that a 2B has nothing to say about llama.cpp issues.

The 9B is the mirror image. It ranks about as well as my 4B did in Part 4, 0.608 against 0.613. It said "yes" 59 times where Strands said it once.

Watch how that number moved while the run was going. After 64 issues it was 0.461, worse than a coin. After 318 it was 0.617. It finished at 0.608. Small samples lie in both directions.

Now the surprise. RSI-Jev is a 2B on the same base model as Strands. It ranked as well as the 9B. The difference between them is 0.007, with a bootstrap interval from -0.057 to +0.038. And it had the best calibration of the lot.

It also said "yes" only twice in 779. So it's cautious, like Strands, but its cautious numbers still put the right issues higher. Strands' don't. Same base, same size, opposite results. The base model isn't the decision model.

Sol landed in between. Better than Strands, and that one's real: +0.069, interval +0.018 to +0.122. Not separable from the rest.

K2-Type is the fastest thing here and the most eager. It said "yes" 356 times. Its average was 48% for a 24% event. Its ranking is barely off a coin.

Gero-4B I stopped at 199 issues. At over three seconds an issue the full run was most of an hour of a hot laptop, and I wanted to know if it was worth that. So I re-scored every model on the same 199 issues:

  Same first 199 issues         AUC   95% interval    Brier
  ──────────────────────────────────────────────────────────
  Clef-flash 9B               0.608   0.513 to 0.703  0.177
  Sol 2B                      0.565   0.467 to 0.661  0.175
  RSI-Jev 2B                  0.562   0.459 to 0.662  0.168
  Gero-4B                     0.544   0.448 to 0.637  0.246
  K2-Type 0.9B                0.541   0.444 to 0.634  0.241
  Strands Decider 2B          0.536   0.438 to 0.630  0.173

Gero sits with K2 and Strands. It over-predicts like K2: 49% on average, against 22% that actually closed in those 199.

But look at the intervals. On 199 issues each one is about 0.19 wide, and every one overlaps every other. That sample can't rank anything. Note RSI-Jev, too: level with the 9B on all 779, a step behind it on the first 199.

It was enough to decide Gero wasn't worth another half hour. It isn't enough to publish a ranking, so I'm not. And if you cut a run short, re-score everyone on the same rows. Comparing Gero's 199 against everyone else's 779 would be comparing different exams.

One smaller test. Six cases from my own work, with three possible actions. K2-Type and RSI-Jev got 4 of 6. Strands, Clef-flash and Gero got 2. Sol got 1. Strands was unsure every time, never above 0.48. Clef-flash picked the same action all six times, with confidence up to 0.93. Gero did the same thing: one action, six times, at 0.76 to 0.90.

Six cases prove nothing. But they show the failure you should fear more. Unsure and wrong gets routed to a human. Confident and wrong doesn't.

On my emails, the easy routing questions, the models agreed nearly everywhere. Five of the six picked the right customer 52 times out of 52. Gero got 46. And RSI-Jev, the best ranker on issues, was the worst on the second routing question, 48 of 52 where the rest got 51.

Which is the other lesson. If the 2B and the 9B agree on your task, you're paying for the 9B out of habit.

The Scores Have the Same Problem

It's not just latency. Every headline number has a footnote doing the real work.

  • The Cloudflare post shows 10 benchmarks. The model card has 41. Jev wins 11 of them, and not narrowly. GPQA Diamond: 78.3 against Clef's 48.0.
  • Clef-flash, the 9B, wins more rows than the 27B does. It also scores 66.8 where the 27B scores 97.4 on one intent benchmark.
  • "64k context window" on the post. max_length defaults to 16,384 on the card.
  • One model in the latency table has a 5.8 ms median and a 222.5 ms p95. A median can hide almost anything.
  • Strands says answers at 0.9 confidence or more are right about 95% of the time. True. That band is 23% of answers. Another 43% come in under 0.5 and are right less than half the time.

Their results page also has the most useful sentence in either release. On pick-one questions, tasks like the ones in training score around 0.96. Tasks it has never seen score around 0.73. Then: "Your own task will land between those two columns."

On a yes-or-no question about a repo it has never seen, Strands landed at a coin flip. RSI-Jev, on the same base, landed level with the 9B.

How They Were Trained

Neither team trained a model from scratch. Both took a Qwen base model, froze it, and trained two small things: an adapter inside the backbone, and a head on the end.

                       Strands Decider      Clef-flash
  ──────────────────────────────────────────────────────
  Base model           Qwen3.5-2B-Base      Qwen3.5-9B
  Adapter (LoRA)       rank 16              rank 256
  Head                 1.05M weights        122M weights
  Adapter file         67 MB                not in the release notes

Rank is the adapter's width. Rank 16 on a 2B is a light touch: the whole fine-tune is a 67 MB file. Rank 256 on a 9B or a 27B is a different size of job.

Strands retrains in about 11 hours on one RTX 3090, using 12.4 GB. That's the honest, small-team number, and it's a good one.

Then open provenance.json in the weights they actually shipped:

"gpu": "NVIDIA H100 80GB HBM3",
"host_shape": "p5.48xlarge",
"train_wall_s": 1685,

Eight H100s. 28 minutes of training, 1 hour 10 for the whole pipeline. The noise study, six retrains to see how much the score wobbles, ran on H100s and A100s.

You can serve this on a small card. You can even train it on one. Finding out whether your change helped, six times over, is still a big-card job.

Clef's answer for fine-tuning is a Cloudflare service. For a 27B with rank-256 adapters, I'd guess that's the only practical answer.

Not Just Those Two

Clef and Strands got the headlines. They're two of about a dozen names I counted in a few weeks, most of which I haven't verified beyond the name. What's interesting isn't the leaderboard. It's that people are making very different bets.

Here are the ones I looked at. What I ran is marked. Everything else is the author's claim, not my finding.

Bet 1: train nothing. SemIf-OpenJev is what I've been running since Part 1. A frozen Qwen3.5-4B, about 1,400 lines of Python, and a softmax over letter tokens. No training at all. It's the honest baseline, and in Part 4 it ranked issues better than anything I built on top of it.

Bet 2: small base, tiny head. Strands is here. So is K2-Type-0.9B, under a billion weights and 2.2 GB, which claims 176 of 231 on the shared benchmark and a 27 ms median. On an H200. And RSI-Jev, on the same 2B base as Strands. I ran both. K2-Type is the fastest thing I tested and over-predicts two to one. RSI-Jev ranked issues as well as the 9B.

Late to the list: Decision 2.0 from vllm-sr, a collection of six models. I ran one, Sol, a 2B. It ships as a Python package you load with trust_remote_code. A 25-line wrapper made it answer the same requests as the others.

Bet 3: big base, heavy head. Clef. 27B, rank-256 adapters, a 122M-weight head on the 9B. The bet is that judgement comes from the size of the model underneath.

Bet 4: train it on a laptop. Gero-4B was fine-tuned on a MacBook, adapting only the last 8 layers, in three stages ending with reinforcement learning for calibration. It has no published benchmarks and doesn't speak the same request format. I ran it anyway, through a small adapter built from the model card's own code, so I'm in the marked group for this one. Its author's write-up is still worth your time, and I'll come back to it.

Bet 5: leave the weights alone, fix the question. jev-align doesn't train anything. It runs your question over your data, finds the rows the model is least sure about, asks you to label those, and rewrites the wording of the question and its options. You approve the diff. Two cautions before you point it at real data: the model doing the rewriting sees your labelled rows, and one of its options publishes them to a public registry.

Bet 6: fit anything anywhere. AirLLM runs enormous models on tiny cards by loading one layer from disk, running it, and throwing it away. The README claims a 70B model in 4 GB. It gives no speed figures. Do the sum from earlier: a 70B model at 16 bits is 140 GB, and every forward pass reads all of it off the disk. Memory was never the only budget.

Bet 7: make more data. If labels are the bottleneck, generate them. I tested this one.

Does Synthetic Data Help?

I took the Part 4 probe and added 500 generated issues to its training set. Two generators: Adaption Labs' zero-seed service, and my own 4B writing issues to a schema. Same 779 real test issues.

  Probe trained on             AUC   Brier     ECE
  ────────────────────────────────────────────────
  Real issues only           0.537   0.184   0.039
  Real + 500 vendor rows     0.535   0.183   0.011
  Real + 500 home-made       0.543   0.184   0.009
  Vendor rows only           0.508   0.252   0.214
  Home-made rows only        0.549   0.197   0.105

No gain in ranking from either. The bootstrap on the difference straddles zero both times.

The calibration column did improve, and I don't trust it. The extra rows pulled the average prediction down from 27% to 25%, toward a test window where 24% of issues close. That's a base rate moving, not a model learning.

Two numbers explain a lot. The vendor rows were 15% "closed fast". The real training data was 35%. And 69 of the vendor's 76 positive rows were compile bugs. Synthetic data that doesn't match your base rate, or hands the model one easy tell, is teaching a different problem.

One test, one task, 500 rows. But the claim on the tin was better data, and on my data it wasn't.

What the Laptop Trainer Learned

The Gero-4B write-up is a list of ways to fool yourself while training one of these. These are the author's findings and numbers, not mine. I'm repeating them because every one applies to grading a model too.

  • A head with 256 answer slots trained six of them. The data rarely had more than six options, so the other 250 stayed random. The fix was one shared scorer applied to each option separately.
  • The model learned "see 'not', flip the answer". Negation words only appeared in negated questions. On a real task, accuracy fell from 0.775 to 0.617. The fix: put negation words everywhere.
  • Wrong answers that weren't in the text gave the game away. "Pick the one that's mentioned" solved the task without reading.
  • A plausible reward destroyed the model. Rewarding "was your confidence in your pick right?" pushed the correct answer's probability from 0.418 to 0.001. Only a proper scoring rule lands on the truth.
  • Brier barely punishes overconfidence. At 0.999 confidence its correction is about 250 times weaker than log score's. Worth knowing, given Clef trains with it.
  • A calibration metric can be rigged by accident. Scored the wrong way, an honest 0.55 looked badly calibrated and an overconfident 0.99 looked perfect.

The last one is the same trap as my 2B. Constant 70% on a 70% task is perfectly calibrated and tells you nothing. Report how well it separates cases, not just how honest its average is.

How People Are Thinking

Put the bets side by side and a few patterns show up.

  • Nobody trains from scratch. Every one I looked at is a Qwen with something small on the end. The base model is the commodity. The head, the data and the loss are the product.
  • The readout is where the ideas are. Letter softmax, pointer head, joint head, one scorer per option. Four designs already.
  • Calibration is a headline number now. Most of the cards I read print ECE or Brier next to accuracy. That's progress, and it's also a new number to game.
  • The speed number still comes from a big card. K2-Type is 2.2 GB and quotes an H200. Same footnote, smaller model.
  • The scores don't transfer. On my easy email task, a 2B and a 9B agreed nearly everywhere. On my hard prediction task, one 2B was a coin flip and another, on the same base, matched the 9B. Neither result is on anyone's leaderboard.
  • The cheapest lever may be the wording. One of these projects changes no weights at all. I haven't tested it. I suspect it's underrated.

Five Posts of Being Wrong

Most of what I learned in this series came from being wrong first. The short version:

  • I thought an idle GPU meant I should batch. A decision model only reads, and batching speeds up writing. No gain.
  • I thought caching the shared part of a prompt was free. This model family carries a running state that can't be rewound. Saving it was 27% of every decision.
  • I thought the model was too big for my Mac. A library default was reserving 256 parallel slots. 12.8 GB of nothing.
  • I thought Python was the slow part. Under 1% of the time. The model's forward pass was 72%.
  • I thought a probe would beat asking. Asking won.
  • I thought a low calibration error meant a good model. It can mean a model that says the base rate to everything.
  • I thought my coverage guarantee held. I'd calibrated and set the threshold on the same rows. The coverage had looked better than it was.
  • I thought more data would help. 500 synthetic rows, no gain.
  • I thought squeezing a model would wreck its probabilities. Done carefully, it moves them by a third of a point.
  • I thought the base model set the ceiling. Two 2Bs on the same base: one a coin flip, one level with a 9B.
  • I thought the model card was the model. Then I ran it.

What Actually Fits

What do most teams want? One box. A T4, or something like it. Sometimes not even that. Sometimes it's whatever CPU is already in the rack.

So here's the ladder, with what's measured and what isn't.

  • CPU. A 2B or a 4-bit 4B, in 3 to 5 seconds per decision. Measured, on a laptop. Fine for a batch. Not for a text box.
  • A laptop GPU. A sub-1B model in a tenth of a second on short text. The 2Bs in a quarter to half a second on short text, 1 to 3 seconds on long. The 9B in 1.4 to 5 seconds, with the fans on. Measured.
  • An 8 GB card. The 9B at 5 bits, half a second for 600 tokens. Measured, by Livesport.
  • A T4. Should hold everything the 8 GB card does, with room. Nobody has published the number.
  • A 24 GB card. The 9B as published. This is where "115 ms" lives.
  • An 80 GB card or bigger. The 27B, and the 38.8 ms.

The models are getting better at the top of that ladder. Most of us are standing at the bottom.

What I Haven't Measured

  • Of the wider wave I've run six. jev-align and AirLLM I've only read. The other five Decision 2.0 models I haven't run. Their numbers are their authors'.
  • The synthetic-data test is one task, 500 rows, and a probe rather than a fine-tune.
  • No T4 run. Not theirs, not mine. The closest is Livesport's 8 GB card.
  • I haven't run Clef, the 27B. It doesn't fit anything I own.
  • Gero-4B's issue run stopped at 199 of 779, by choice. On those 199 it can't be told apart from any other model.
  • The laptop table compares software stacks and number formats as well as models. I haven't separated them.
  • Every model on the issues is zero-shot with one wording of the question. A different wording could move them. I haven't investigated why Strands came out at a coin flip and RSI-Jev, on the same base, didn't.
  • Six triage cases prove nothing on their own. One case moves the score by 17 points.
  • The Clef latency rows may not be from the H200. The release doesn't say.
  • The T4 token counts use a peak figure and a rough FLOP count.
  • The card specs are from vendor sheets, and the quantization table is Livesport's measurement, not mine.
  • My email set is small and private, so I'm giving you its latency and not much else.

Still owed: Strands 2B on a real T4, cold and warm, p95 not median.

The Point

In Part 2 I said to measure on the card you ship on.

This week I learned the other half. Read the benchmark on the card it was run on.

You don't need the vendor to tell you. Count the weights. Multiply by the bits. Count your tokens. That's memory and time, on any machine, before you download a byte.

The question isn't how fast the model is. It's whose GPU.

Read These

In Plain English

  • Two new AI models came out that make quick decisions instead of writing text. Both advertise very fast answers.
  • Those speeds were measured on big, expensive graphics cards. Most teams have a small one, or none.
  • You can work out whether a model fits your machine with one sum: how many numbers are in it, times how many bits each one takes.
  • How long it takes depends on how much text you give it. Twice the text, about twice the wait.
  • I ran six of them on my laptop. The smallest answered in a tenth of a second on short emails. The two-billion ones took a quarter to half a second, and one to three seconds on longer text. The biggest took one to five seconds and ran the laptop hot.
  • With no graphics card at all, expect three to five seconds per answer. Too slow for a live screen, fine for an overnight job.
  • Shrinking a model to fit a small card barely changes its answers if it's done carefully. Done carelessly, it does.
  • Fast doesn't mean right. On my own test, one small model was no better than a coin flip at picking which issues would close. Another the same size, built on the same base, picked as well as the biggest. The fastest one said "yes" far too often.
  • Lots of other small decision models appeared in the same few weeks. They all start from the same free base model and add a small piece on top. Almost all quote speeds from a big card.
  • I tried adding computer-generated examples to my training data. It didn't help.
  • The lesson: when you see a speed number, ask what machine it was measured on. If nobody says, assume it's bigger than yours.