Pavan Kumar T V

CTO | Technology Leader

By ·

Hold My Beer

Series: start with Part 1 on what a decision model like Jev is good for, Part 2 on what it does to a GPU, then Part 3 on why deciding is not predicting. This is Part 4, where I build the predictor.

Yes, a post called Hold My Beer, published on 2 October. Gandhi Jayanti. A dry day. My Indian timeline is already typing "blasphemous". Relax, the beer is a metaphor. Nobody drinks in this one.

When Jev blew up, the whole timeline learned the same trick. Don't make the model write. Make it decide. Ask a question, read one forward pass, done.

It's a great trick. I've been running an open version of it for days.

But watching everyone reach for "just ask the model" for every problem, I had a very old reflex. The one you get from years of shipping forecasts that were wrong in production.

Hold my beer.

I took the same 4B model and refused to ask it anything at all.

Words You'll Need

Part 3 covers hidden states, probes, calibration and conformal prediction. These are the scorecard and the rest of the toolbox.

  • Zero-shot. Asking the model straight away, with no training on your data. The "just ask it" approach.
  • Baseline. The dumbest model that could work. If the clever one can't beat it, the clever one isn't clever.
  • Base rate. How often the outcome happens overall. Always guessing it is the first baseline.
  • Time split. Train on older examples, test on newer ones. The way the model will be used for real.
  • Logistic regression. A plain model that adds up weighted inputs and turns the total into a probability. Old, fast, hard to fool yourself with.
  • Cross-validation. Trying settings on several held-out slices of the training data and keeping the one that holds up best.
  • Regularization. A penalty that stops a model leaning too hard on any one input. How strong it is gets picked by cross-validation.
  • Drift. The world changing under your data, so last spring's patterns stop holding.
  • AUC. Pick one issue that closed fast and one that didn't. AUC is how often the model ranks the fast one higher. 0.5 is a coin flip, 1.0 is perfect.
  • Brier score. The average squared gap between the predicted probability and what happened. Lower is better.
  • ECE. Expected calibration error. Group predictions by confidence and measure how far "said 30%" sits from "happened 30% of the time". Lower is better.

The Bet

Last post I argued that deciding and predicting are different jobs. A decision is a judgement about evidence you already have. A prediction is a claim about a future the record doesn't contain, and the only honest teacher for that is history.

I also admitted I hadn't built anything. So I built it.

The bet: keep the model, throw away the question. Use the LLM only to read. Hand what it read to the boring ML toolbox from twenty years ago. Baselines, time splits, linear probes, calibration, conformal prediction. Then put the two approaches side by side on the same data and see who wins.

The Data

I can't show you work data, so I picked something public that has the same shape as my real problem: messy text written by humans, a few structured fields, and an outcome that only arrives later.

GitHub issues on llama.cpp. The question: will this issue be closed within 7 days of being opened?

  • 3,113 issues, opened between January and early September 2026
  • Each one frozen as it looked the moment it was opened. Title, body, who opened it. No comments, no labels added later, no close date
  • Label: closed within 7 days or not. Only issues at least 30 days old, so every label is settled

And one thing I didn't plan for. The share of issues closed within a week fell from about 42% in the first quarter to about 24% by the summer. The world changed under the data. That turned out to be the most useful thing in the dataset.

The Toolbox

Here's the thing: none of what follows is new. That's the point. Every technique below has decades of practice behind it, and almost none of it shows up when people put LLMs into production.

  ASK  (the decision readout)

    issue ──▶ 4B model ──▶ "yes" / "no" letter logits
                           probabilities from vibes


  PROBE  (old ML on a frozen model)

    issue ──▶ 4B model, one prefill ──▶ 2,560 floats
                                           │
    numeric fields ────────────────────────┤
                                           ▼
                           logistic regression, tuned
                           with time-ordered folds
                                           │
                           calibrated on the recent tail
                                           │
                           conformal threshold
                                           ▼
                    P(closed in 7 days) + a set with a
                    coverage guarantee

1. Baselines first

Before any model, ask what the dumbest forecast scores. Predict the training base rate for every issue.

On the last 779 issues, that scores a Brier of 0.195 (mean squared error of the probability, lower is better). It's also badly miscalibrated, with an expected calibration error of 0.108, because the base rate it learned is from a world that no longer exists.

Then a model with only cheap structured features. Title and body length, code blocks, whether there's a stack trace, the issue template type, day and hour, whether the author is a contributor.

AUC 0.55. Barely better than a coin at ranking issues. Brier 0.183.

That's the bar. Anything with an LLM in it has to beat a number most people never compute.

2. Split by time, never by shuffle

Every split in this project is chronological. Train on January to mid-July. Test on mid-July to September.

A random split mixes weeks. The model sees issues from the same week in training and test, learns what that week was like, and looks better than it is.

I ran the numeric model both ways. Same rows, same code, only the order changed.

  Split       AUC
  ───────────────
  Random    0.584
  By time   0.550

Not a dramatic gap. That's what makes it dangerous. It's small enough to believe, and it's pointing the wrong way. The random split tells you the model is getting better when the world it will run in is getting harder.

3. A probe instead of a prompt

The model reads each issue once. No question, no options, no answer. I keep the final hidden state of the last token: 2,560 numbers that the output layer would normally squash into a letter.

A logistic regression learns from those numbers which issues get closed fast. That's a linear probe. Interpretability researchers use them to ask what a model knows. I'm using one to ask what it's worth.

Not much, it turns out.

  Model                   AUC   Brier     ECE
  ───────────────────────────────────────────
  Numeric fields only   0.550   0.183   0.023
  Probe                 0.537   0.184   0.039
  Probe + numeric       0.539   0.184   0.037

The probe on 2,560 numbers from a 4B model did no better than a handful of cheap fields. A bootstrap on the difference runs from the probe being 0.05 better to 0.08 worse. In plain words: no real difference, and adding the fields to the probe didn't help either.

4. Tune on the past, test on the future

Regularization strength is picked by five-fold cross-validation where every fold trains on earlier data and validates on later data. Scikit-learn calls it TimeSeriesSplit. It costs one argument.

5. Calibrate on the most recent slice

The last 20% of training (the newest data) never touches the model. Half of it fits a two-parameter sigmoid that maps raw scores to probabilities. Half sets the conformal threshold below.

That alone took the numeric model from the base rate's 0.108 calibration error down to 0.023. The calibrator learned that the world had moved, because it only saw the recent world.

That's drift correction for free, and it's two lines of code.

6. Conformal prediction

A probability is a claim. Conformal prediction turns it into a promise.

From the held-out slice, pick a threshold so that 90% of the time, the true answer is inside the set you return. If the model is sure, the set is one label. If it isn't, the set is both labels, and that's the model honestly saying "I don't know".

For the numeric model: 87.8% coverage against a 90% target, with an average set size of 1.46. Slightly under, which is what drift does to a guarantee that assumes tomorrow looks like yesterday.

The probe hit 90.2% coverage, with an average set size of 1.57. On target, but with bigger sets. It kept its promise mostly by saying "I don't know" more often.

7. Cache the expensive part

The forward pass is the only expensive step. Every vector is cached by a hash of model, settings and text. Ten different targets on the same issues cost one pass, not ten. Retraining the head takes seconds on a CPU.

Ask vs Probe

Now the head-to-head. Same 779 test issues.

                      AUC   Brier     ECE   Coverage
  ──────────────────────────────────────────────────
  Base rate         0.500   0.195   0.108
  Numeric fields    0.550   0.183   0.023      87.8%
  Ask (zero-shot)   0.613   0.190   0.107
  Probe             0.537   0.184   0.039      90.2%
  Probe + numeric   0.539   0.184   0.037      89.9%

The probe lost. Asking won.

Asking the model ranked issues better than anything else on the table. The bootstrap puts its lead over the probe between 0.01 and 0.14 AUC, and over the numeric fields between 0.002 and 0.12. Clear over the probe, only just clear over the fields.

But look at the other two columns. Its probabilities were bad. On average it gave a 14% chance of a fast close, when the real rate was 24%. It said "yes" outright on fewer than 1 in 100 issues. Its calibration error was the same as blindly guessing the base rate.

So the model knew which issues were more likely to close. It just didn't know how likely.

That's a problem the toolbox has solved for decades. So I stopped racing the two approaches against each other and stacked them. I calibrated the model's answer on the first half of the test window (mid-July to mid-August) and scored the second half, 390 issues it had never seen.

  Second half              AUC   Brier     ECE
  ────────────────────────────────────────────
  Numeric fields         0.537   0.182   0.036
  Probe                  0.532   0.183   0.041
  Ask                    0.625   0.186   0.108
  Ask + calibration      0.625   0.176   0.024
  Ask + numeric fields   0.632   0.176   0.015

Same ranking, now with honest probabilities. Calibration error down from 0.108 to 0.024, and the best Brier on the board. Add the cheap fields and it gets a little better again.

The winner wasn't old ML instead of the LLM. It was old ML wrapped around the LLM's answer.

The GPU Bit

Of course there was a GPU detour. There always is.

The first time I asked llama.cpp for embeddings on my Mac, it died with an out-of-memory error. Even on "hello world". The normal scoring path, same model, same machine, was fine.

The verbose log had the answer. In embedding mode, the Python bindings reserve room for 256 sequences at once, so you can embed a batch in parallel. On a normal transformer that's cheap. On this hybrid model, every sequence carries its own recurrent state, about 50 MB. 256 of them is 12.8 GB, which is right at the limit of what Metal will give a single process.

The fix was to skip embedding mode entirely. Open a normal single-sequence context, switch on embeddings with one llama.cpp call, read the last token's vector. The vectors match the CPU path at a cosine similarity of 0.996 or better, and on my probe it ran in 0.8 seconds per issue instead of 15.

Same lesson as last time. Read the log before you guess.

What This Doesn't Prove

  • One repository, one target, one horizon. A different question could flip the result.
  • "Who opened it" is the author's relationship to the repo today, not on the day they opened it. A small leak, and it only helps the models that use the numeric fields.
  • Some fast closes are duplicates and spam. A model that spots those is useful, but it's not predicting engineering effort.
  • The ask-plus-calibration result is one half-window of 390 issues. It's a direction, not a verdict.
  • I haven't tested why the probe failed. Too few rows for 2,560 numbers, the last token instead of an average over all of them, or a 4B model that just doesn't hold this in a straight line. Each of those is an experiment, not a conclusion.

The Point

I said hold my beer and refused to ask the model anything.

I spilled it. The question was the best feature I had.

But the question alone was a confident liar about how sure it was, and nobody would have noticed without a baseline, a time split and a calibration check. Twenty-year-old tools caught it. Two more lines fixed it.

Don't pick between the LLM and the toolbox. Ask the model. Then grade it like it's 2006.