Pavan Kumar T V

CTO | Technology Leader

By ·

Wrong Tree

Series: Part 1 is what a decision model like Jev is good for. This is Part 2, on what it does to a GPU. Part 3 asks why deciding is not predicting, and Part 4 builds the predictor and races it against asking the model.

I had a decision model labelling email on a T4 at work. Most of a second per email.

My first instinct was the same as yours. The GPU must be sitting idle. Python glue, one email at a time, state bouncing through host memory. Batch it, keep everything on the card, done.

So I had an agent build exactly that. GPU-resident prefix copies. Eight emails per decode call. A clean PR with tests.

It measured flat. 761 ms per request serial. 741 ms batched.

I stared at those two numbers for a while. Twenty milliseconds. That was the entire reward for all that clever plumbing.

Here's the thing: the GPU was never idle. My mental model of the workload was wrong. And it was wrong in a way that most "make inference faster" advice is wrong for this whole class of model.

This post is the critical review I wish I had done before writing a line of code. Full GPU detail, with every number labelled by where it came from.

Words You'll Need

If you've never run a model on a GPU, skim this first. If you have, skip ahead.

  • Token. A chunk of text, roughly three quarters of a word. Models read and write tokens, not characters.
  • Forward pass. One trip through every layer of the model. The basic unit of work.
  • Prefill. Reading the prompt. All the prompt tokens go through together, side by side, in one big forward pass.
  • Decode. Writing the answer, one token per forward pass. A 200-token reply is 200 trips.
  • Logits. The raw score the model gives every token in its vocabulary for "what comes next". Softmax turns those scores into probabilities.
  • KV cache. The model's notes on tokens it has already read, so it doesn't reread the whole prompt for every new token.
  • Quantization. Storing the weights in fewer bits. Q4_K_M is a llama.cpp format at roughly 4 to 5 bits per weight. Smaller and faster to read, slightly less precise.
  • GGUF and llama.cpp. The file format, and the C++ engine that runs these quantized models on almost anything, from a MacBook to a datacenter card.
  • FLOPs. Floating-point operations: the multiplies and adds. TFLOPS is trillions of them per second, which is how fast a chip does math.
  • Memory bandwidth. How fast a chip can read its own memory, in GB/s. Every forward pass has to read the weights, so this is often the real limit.
  • T4. A small, older Nvidia datacenter GPU with 16 GB. Cheap, and available on every cloud.

A Decision Model Never Generates a Token

I wrote about Jev-style decision models last week. Three kinds of question: pick one, put it on a scale, yes or no.

The open reproduction I run is SemIf, on Qwen3.5-4B, quantized to Q4_K_M, through llama.cpp. The trick is in the readout. It never asks the model to write an answer. It builds a prompt that ends exactly where the answer letter would go, runs one forward pass, and reads the logits of the option letters. In other words: "if you were to say A, B or C right now, how much would you want to say each one?" Then it never lets the model say anything.

  one prompt, one forward pass

  ┌────────┬────────────────────┬───────────┬─────────────┐
  │ system │ evidence           │ criterion │ A: …  B: …  │
  │ fixed  │ 80 to 1,100 tokens │ fixed     │ fixed       │
  │        │ varies per email   │           │             │
  └────────┴────────────────────┴───────────┴──────┬──────┘
                                                   │
                                    last position only
                                                   ▼
                          logit[A]  logit[B]  …  logit[P]
                                                   │
                                                softmax
                                                   ▼
                 { choice: "B", probabilities, confidence }

  no sampling loop · no JSON parsing · no second token

SemIf's own README puts this at 1.02 s against 5.33 s for generating the same answers as JSON on a 3090. That gap is the whole product.

But notice what it does to the workload. There is no decode phase. Every token the GPU touches is prefill. A chatbot reads your question and then writes for a while. A decision model only reads. Hold that thought, because it decides everything below.

What's Actually Inside the Model

I read the GGUF metadata instead of trusting what I remembered. Good thing, because what I remembered was wrong.

  • 32 transformer blocks, plus one multi-token-prediction block (an extra head for guessing several tokens ahead, only useful when generating) that a decision never uses
  • full_attention_interval = 4. Only 8 of the 32 blocks are real attention, the kind that can look back at every earlier token. The other 24 are Gated DeltaNet, a linear-attention recurrent layer. Instead of a transcript of every token, it keeps one running summary of fixed size. A tally, not a transcript
  • Hidden size 2,560. Attention: 16 query heads, 4 KV heads, head dimension 256
  • Vocabulary: 248,320 tokens. I had been quoting 151k from older Qwen models. Wrong by 100k rows

That hybrid layout matters more than anything else on this list.

For the 8 attention layers, the cache grows with length. 8 layers × K and V × 4 heads × 256 dims × 2 bytes is 32 KB per token. A 1,000-token email is about 32 MB.

For the 24 DeltaNet layers, there is no per-token cache at all. Each layer keeps one fixed state matrix, 32 heads × 128 × 128 in fp32, about 2 MB per layer. Roughly 48 MB for the stack, whether the email is 50 tokens or 50,000. (Arithmetic from the metadata, not measured.)

And here's the catch: you cannot rewind a recurrent state.

An attention cache is a list. Want the state after the evidence but before the question? Truncate the list. A DeltaNet state is a running sum. Every token has already been folded into it. There's no subtracting the question back out. llama.cpp knows this and refuses seq_rm (its "forget the last N tokens" call) on the tail of a hybrid sequence.

So to ask two questions of one email, you either save a snapshot before the first question and restore it before the second, or you copy the sequence.

Two Ways to Ask Twice

  SERIAL  save and restore through host RAM

     GPU                         host RAM
   ┌──────────────┐  save     ┌──────────┐
   │ prefill      │──────────▶│ snapshot │
   └──────────────┘           └────┬─────┘
   ┌──────────────┐  restore       │
   │ + question 1 │◀───────────────┤
   └──────────────┘                │
   ┌──────────────┐  restore       │
   │ + question 2 │◀───────────────┘
   └──────────────┘
   N questions = N restores, one decode each


  FAN-OUT  llama_memory_seq_cp, never leaves the GPU

   ┌──────────────┐
   │ prefill      │ seq 0
   └──────┬───────┘
          │ seq_cp: N branches share the prefix cells
     ┌────┴─────┬──────────┬──────────┐
     ▼          ▼          ▼          ▼
   ┌────┐     ┌────┐     ┌────┐     ┌────┐
   │ q1 │     │ q2 │     │ q3 │  …  │ qN │
   └─┬──┘     └─┬──┘     └─┬──┘     └─┬──┘
     └──────────┴────┬─────┴──────────┘
                     ▼
         one llama_decode, N logit rows

On my M3 Pro I profiled the serial path: 72% of scoring time in llama_decode, 27% in save_state at 207 ms per email, and under 1% in Python.

Two conclusions fall out of that.

First, rewriting the glue in Rust buys you nothing. The glue is already under 1%.

Second, the fan-out kills the 27%. Upstream SemIf PR #34 does exactly this and measured 13 to 18 percent on a 3080. I merged it into my fork.

Now the critical part. That 207 ms is a Mac number. Apple Silicon has unified memory, so it was never a bus transfer. It's serialization overhead. On a T4 the same ~80 MB crosses PCIe Gen3, the link between the CPU's memory and the GPU's, which is single-digit milliseconds per copy at realistic rates. The snapshot cost may be big on the Mac and small on the T4. I optimized the platform I had, not the one I ship on.

The Roofline Says Batching Is a Decode Trick

So why didn't batching eight emails together help?

Think of a kitchen. Memory bandwidth is how fast ingredients arrive from the store room. Compute is how fast the chef chops. If the chef spends the day waiting on deliveries, cooking eight dishes from one delivery is a huge win. If the chef is already chopping flat out, batching orders changes nothing. The chef was never idle.

A roofline is that picture with numbers on it.

Back-of-envelope, labelled as such. The T4 has about 65 TFLOPS of fp16 tensor-core compute and 320 GB/s of memory bandwidth. Divide one by the other and you get the ridge point: about 200 FLOPs per byte. Below that you wait on memory. Above it you wait on math.

Qwen3.5-4B has roughly 3.5B non-embedding parameters, about 2.5 GB at Q4_K_M. Each token costs about 2 FLOPs per parameter.

  • Decode at batch 1: one token per weight read. About 3 FLOPs per byte. Deep in memory-bound territory. This is where batching works, because eight sequences share one weight read.
  • Prefill, 512-token micro-batch: 512 tokens per weight read. About 1,400 FLOPs per byte. Seven times past the ridge. Already compute-bound before you batch anything.
  FLOP/s
   65T ┤            ╭━━━━━━━━━━━━━━━━━━━━━━━━●━━━━  compute roof
       │          ╱ ┆                       prefill, 512 tokens
       │        ╱   ┆                       ~1,400 FLOP/byte
       │      ╱     ┆                       ← decision models
       │    ╱       ┆
       │  ╱         ┆
       │●           ┆
       └────────────┴────────────────────────┴────▶ FLOP per
       1           200                    1,400     byte read

   ● bottom left  decode, batch 1, ~3 FLOP/byte. Chatbots live
                  here; batching climbs the 320 GB/s slope.
   ┆ ridge        65 TFLOPS ÷ 320 GB/s ≈ 200 FLOP/byte.
   ● top right    prefill is already on the compute roof.
                  Batching has nothing left to amortize.

A chatbot spends most of its life at the bottom-left dot. Every serving trick you've read about (continuous batching, paged attention, speculative decoding) exists to drag that dot up the slope.

A decision model lives at the top-right dot. It's already pinned against the compute roof. Batching moves nothing because there's nothing to amortize.

My M3 measurement agrees: serial and batched both ran at about 2.1 ms per token. Same roof, same speed.

The same arithmetic gives a floor. A 1,000-token prompt is about 7 TFLOPs. At the T4's peak that's around 110 ms. Real quantized kernels run well below peak, so "most of a second per email" with two questions is not a GPU sitting idle. It's a GPU doing exactly the math I asked for.

The only lever on a compute-bound workload is to ask for less math.

The Critical Review

Here's what I'd score as right, wrong, and still unknown.

Right

  • Logit readout. No decode is the single biggest win, and it's free.
  • GGUF on Turing. Turing is the T4's chip generation, and it has no native bf16, the 16-bit number format newer cards handle in hardware and SemIf's Torch path uses. Running the Torch bf16 path would mean fighting the hardware. The quantized path sidesteps that, and the job log showed all layers offloaded to the GPU with flash attention (a faster, memory-friendly way to compute attention) on.
  • Fail loud. The first version of the fan-out quietly fell back to the serial path on any error. Errors became slowdowns nobody saw. It now raises. A slow wrong path is worse than a fast crash.

Wrong

  • The idle-GPU hypothesis. Disproven by my own measurement. Should have done the roofline first. It takes ten minutes.
  • Slicing the output head. I'd suggested computing only the 16 letter rows of the 248,320-row output matrix. It sounds great: 0.006% of the rows. But the head is only read for the last position of each question. On a T4 that's around 1.6 ms of a second-long request. A rounding error dressed up as an optimization.
  • The permutation bomb. SemIf can reduce option-order bias by scoring each question under K shuffled option orders and averaging. Upstream measured 10 of 36 decisions flipping under shuffles, so this is worth having. But the code picked K orders by first listing every permutation. Ten options is 3.6 million orderings. 1.45 seconds of CPU before the GPU saw a token. Sampling them lazily brought 40 options down to 0.05 ms. Fixed in my fork.

Still unknown

  • Every optimization above, on a T4. All the evidence is from a Mac, a 3080, or arithmetic. None of it is from the card in production.
  • Q8_0 against Q4_K_M. K-quant dequantization costs compute, and compute is exactly what we're short of. Q8_0 fits easily in 16 GB and may prefill faster on Turing's tensor cores, the units built for exactly this kind of matrix math. It may also shift answers, so it needs an accuracy rerun alongside.
  • Mid-layer readout. Interpretability work (researchers opening up models to see what each layer is doing) says the answer often settles before the last layer. If a decision commits at layer 22 of 32, that's a third of the FLOPs gone. That's the only idea here that cuts the math itself.

Where the Real Speed Is

If prefill is compute-bound, the cost is tokens × parameters. So you attack tokens or parameters. Nothing else.

Tokens. The prompt puts the email first and the question and options after. The options are identical for every email but sit after the varying part, so they get re-prefilled every time. Put them first and they're computed once per question for the entire run. The catch is the same order sensitivity again: the published accuracy was measured on the current layout, so moving it means re-running the eval. Then there's the body excerpt length, and threads scored once per reply instead of once per conversation. All product decisions, not kernel decisions.

Parameters. Early exit at the layer where the answer commits. A smaller base for easy questions with the 4B as the escalation. Distilling the decisions into an encoder for the hot path.

None of that is a serving trick. All of it is about the shape of the question.

The Lesson

I've been building software for 25 years, and I still walked straight into this. Worse, I had an agent build the wrong thing beautifully, with tests, before I asked whether it was the right thing.

The serving playbook everyone shares was written for chatbots. Chatbots decode. Decision models don't. Same GPU, same model weights, opposite bottleneck.

Do the roofline before you write the PR. Measure on the card you ship on.

The GPU was never idle. I was just asking it the wrong question.