By Pavan Kumar T V ·
Not So Fast
Series: Part 1 is what a decision model like Jev is good for, Part 2 what it does to a GPU. This is Part 3, on why deciding is not predicting. Part 4 builds the predictor and puts it to the test.
My decision model can tell you whether an email is a complaint.
It cannot tell you whether that shipment will be late.
Those sound like the same kind of question. Same model, same evidence, same three answer types. They are not. And the difference is the most useful thing I've figured out about these models since I started running one.
Words You'll Need
Part 2 covers tokens, forward passes and logits. These are the new ones.
- Label. The right answer for one example. For a decision, a person writes it. For a prediction, the world writes it later.
- Hidden state. The list of numbers the model holds for a token after reading it. The last token's hidden state is the model's whole summary of the text, just before it gets squeezed into an answer.
- Residual stream. The running hidden state as it passes up through the layers. Each layer reads it and adds a bit.
- Embedding. A text turned into a fixed list of numbers, so ordinary maths can work with it.
- Linear probe. A small, plain model (one weight per number) trained on top of hidden states. If it works, the information was already in there.
- Calibration. Whether "70% sure" comes true about 70% of the time. A calibrated model's confidence means something.
- Conformal prediction. A way to wrap any model's answer in a range or set that holds the true answer a promised share of the time, say 90%.
- Leakage. When the training data quietly contains the answer, so the model looks brilliant in testing and fails in real life.
- scikit-learn. The standard Python library for classic machine learning. Every piece in this post is a few lines of it.
The Itch
I've written about decision models and about what they do to a GPU. Pick one, put it on a scale, yes or no. One forward pass, read the letter logits, done.
After a few days of that, the obvious next question showed up. If it can score "how urgent is this email" on a 1 to 5 scale, why not "how many days late will this be"? Same readout. Make the scale levels into buckets. On time, one day, two to three days, a week, worse.
You can do that today. SemIf's score type already returns a probability per level and the expected value across them. I checked the code. It's a request away.
And it would be a demo, not a product.
Same Readout, Different Teacher
Here's the thing: the readout isn't what makes something a decision or a prediction. The teacher is.
DECIDE the truth is in the record
email ──▶ model ──▶ "complaint"
▲
│ a human reads the same email
│ and agrees or disagrees
judged by: agreement with a human
metric: accuracy
PREDICT the truth arrives after the record
record, Mon ──▶ model ──▶ "3 days late"
▲
Mon ──── Tue ──── Wed ──── Thu ─┴── parcel arrives
the world labels it
judged by: what actually happened
metric: calibration, interval coverage
A decision is a judgement about evidence you already have. Somebody could read the email and agree or disagree. The answer is in the record.
A prediction is a claim about a future the record doesn't contain. No amount of reading the email tells you whether customs releases it on Wednesday. The only teacher is history: thousands of past records, each paired with what happened next.
That's why the zero-shot score is a toy. A 4B model asked "how late will this be" has never seen your lanes, your carriers, your customs brokers. It will produce a confident-looking distribution from vibes. Its probabilities are, in SemIf's own words, "uncalibrated as decision confidence."
So the question became: what does a proper ML version look like?
I Looked for It. It Doesn't Quite Exist.
Prediction from records with text in them is not new. Three families already cover parts of it.
- Fine-tune a small encoder per target. ModernBERT with a regression head, or embeddings into gradient boosting. Strong, boring, needs a few thousand labels per target.
- LLMs as regressors. OmniPred (2402.14547) trains an LM to regress across tasks from text. LLM Processes (2405.12856) gets full predictive distributions conditioned on language. Slow, and calibration is modest.
- Tabular foundation models. TabPFN (2207.01848) and TabICL (2502.05564) predict a new table with no training by reading the labelled rows in context. Brilliant on numbers. Text columns get dropped or embedded crudely. A benchmark for exactly that gap exists (2507.07829), which tells you the gap is real.
Then I searched for the obvious packaging: an LLM as a scikit-learn estimator. There is a scikit-llm. It's prompt-and-parse behind a fit/predict signature, calling hosted APIs. No logits, no hidden states, no regression, no calibration.
Nobody has the version where the LLM is a real estimator. One you can drop into a Pipeline, read probabilities from, calibrate, cross-validate, and wrap in conformal intervals.
That's the thing I want.
The Model Already Knows More Than It Says
The interpretability literature hands you the design.
- The answer is readable from the residual stream before the last layer (tuned lens, 2303.08112).
- Linear probes on hidden states predict correctness better than the output probability does (2304.13734, 2310.06824).
- Letter readouts carry a prior toward certain letters regardless of content (2309.03882).
- LLM embeddings beat hand-built features for regression, and the gain grows with input complexity (2411.14708).
Put those together and the conclusion is uncomfortable for anyone selling a bigger decision model.
The letter softmax is the worst place to read from. It's the last, narrowest, most biased layer of a network that has already built a rich representation of the record. A decision model throws that representation away and keeps 16 numbers.
For prediction, keep the representation. Throw away the letters.
The Design
record as of the cut-off
┌───────────────────────────┐
│ state: JSON + email text │
│ numeric fields │─────────────┐
└─────────────┬─────────────┘ │
▼ │
┌───────────────────────────┐ │
│ frozen Qwen3.5-4B, GGUF │ │
│ one prefill, no decode │ │
│ last-token hidden state │ │
└─────────────┬─────────────┘ │
│ 2,560 floats │
├──▶ cache by content hash │
▼ ▼
┌─────────────────────────────────────────────────┐
│ scikit-learn head on [ vector | numerics ] │◀── fit on
│ linear probe or gradient boosting │ history:
└────────────┬───────────────────────┬────────────┘ record +
▼ ▼ outcome
calibrated P(late) conformal interval
"2 to 5 days, 90%"
Four pieces, all deliberately boring.
An encoder. A scikit-learn transformer. fit does nothing. transform runs each record through the frozen model once and returns its hidden-state vector. Vectors are cached by content hash, so ten prediction targets on the same record cost one forward pass.
A head. Ridge, logistic regression, or gradient boosting on the vector plus whatever numeric fields you have. It trains in seconds on CPU. Retraining when the world shifts is a cron job, not a GPU rental.
Calibration. For classification, isotonic or sigmoid calibration out of the box. If the model says 30% late, about 30% should be late. SemIf's own temperature fit took one benchmark's calibration error from 0.208 to 0.069, so this step is not decoration.
Conformal intervals. For regression, hold out a slice of history, measure the errors, and use their quantile as the interval width. You get a statement like "2 to 5 days, 90% coverage" that holds without trusting the model's own sense of confidence.
None of this is new ML. That's the point. The novelty is plugging a 4B model into the ecosystem that already solved evaluation. GridSearchCV over prompts and layers. learning_curve to see how many labels you need. Drift tests on the vector distribution. SHAP on the head. A decision model gets none of that today.
The Trap: Time
There's one way to make this look amazing and be worthless.
history ─────────────────────────────────────────────▶ time
RANDOM SPLIT wrong
┌────┬────┬────┬────┬────┬────┬────┬────┬────┬────┐
│ tr │ TE │ tr │ tr │ TE │ tr │ tr │ TE │ tr │ tr │
└────┴────┴────┴────┴────┴────┴────┴────┴────┴────┘
test weeks sit between train weeks,
so the model learns the week, not the shipment
TIME SPLIT right
┌──────────────────────────────────────┬───────────┐
│ train │ test │
└──────────────────────────────────────┴───────────┘
test is strictly after train, like production
and inside every row:
state = the record AS OF the cut-off
strip delivered_at, claim_id, the "arrived, thanks" reply
Two kinds of leakage, both silent.
Leakage across rows. Shuffle and split, and the test set is full of records from the same weeks as the training set. The model learns the week, not the shipment. Split by time or don't bother.
Leakage inside a row. If the record you encode was pulled today, it already contains the answer. The delivery scan. The follow-up email saying "arrived, thanks." The model will find it, because finding things in text is the one thing it's great at. Every record has to be a snapshot from before the outcome.
A decision model never had this problem. Its answer was supposed to be in the record.
What I've Actually Measured
Honestly? Almost nothing yet. One probe.
I loaded the same Q4_K_M GGUF in llama.cpp's embedding mode with last-token pooling. On CPU it returns a 2,560-dimension vector per record. Three toy sentences: two about a customs hold came out closer to each other (cosine 0.834) than either was to an on-time delivery (0.79). That proves the vectors aren't noise. It proves nothing about prediction.
It took 8.7 seconds per record on a cold CPU. On the Mac's GPU, the embedding context ran out of memory, which the regular scoring path on that same Mac never does. I don't know why yet.
What I haven't done is the only experiment that matters.
The Experiment That Decides It
Pick one outcome with history. Late delivery, claim filed, return, time to reply. Build a few hundred to a few thousand rows of record-as-of-cut-off and what happened. Split by time. Then train three models on the same split:
- A dumb baseline. Majority class, or the mean.
- Gradient boosting on the structured fields only. No LLM.
- The same, plus the LLM vector.
If model 3 doesn't beat model 2 on calibration and held-out error, the 4B model is an expensive way to read text you didn't need. Kill it.
If it does, you've got something no decision model offers: one forward pass per record, cached, feeding as many calibrated predictions as you have outcomes to learn from. Decisions and forecasts off the same vector.
I'm building it as an enabler in my SemIf fork. StateEncoder, SemIfPredictor, and an eval command that does the three-way comparison with a time split by default. It isn't merged, it isn't measured, and I'll publish the numbers whichever way they go.
The Point
Everyone in this space is racing on the same axis. Bigger base, better readout, higher benchmark.
But a decision tells you what the evidence says. A prediction tells you what happens next. Operators pay for the second one.
Stop asking the model what it thinks. Ask it what it sees, then let history do the predicting.