What if your model didn’t generate free-form text at all, but instead scored the allowed answers directly and returned calibrated probabilities? That’s the core idea behind Jebadiah v2.1, a pair of open-weight models that treat decisions as a closed-set scoring problem rather than a generation task. At AI Tech Inspire, we spotted this release because it reframes a familiar workflow—question in, options out—in a way that could make evaluation, deployment, and safety checks a whole lot simpler.
Quick facts from the release
- Models: Jebadiah v2.1 in 27B and 9B sizes; open weights,
Apache-2.0license; based on Qwen. - Approach: Closed-set decision modeling. Input = structured context + typed question + fixed options. Output = a probability per option from
candidate-label logits; no sampling or parsing. - Benchmark:
Decision Index 0.3, self-run with the official scorer; submitted to the board (which adds private tests before ranking). - 27B results: 57.03 vs 55.11 (v2). Category deltas: Knowledge +2.34, Language +3.89, Retrieval +2.81, Arts +1.31, Tools -2.10.
- 9B results: 47.20 vs 44.09. Category deltas: Knowledge +3.67, Language +4.34, Retrieval +5.96, Arts +0.01, Tools -0.84.
- Regression hotspot:
When2Callfell 6.38 (27B) and 7.44 (9B); cause unknown; deltas not claimed to be statistically significant. - Contamination check: Public benchmark families overlapping training sources were text-matched with zero hits (paraphrases could still slip through).
- Quantization: Agreement with full weights measured on 260 held-out questions per build. Examples: 27B
Q8_0259/260, 9BMLX 4-bit237/260; per-question records published. - Resources: Weights on Hugging Face — 27B: frontier-infra/jebadiah-27b-v2-1, 9B: frontier-infra/jebadiah-9b-v2-1. Results: decision-index-results. Code and evals: github.com/getainode/jebadiah.
What’s different: scoring labels from logits
Most teams reach for a GPT-style generator and then coax it into multiple choice by string formatting and careful prompting. Jebadiah flips that script. It takes a fixed option set and computes probabilities straight from candidate-label logits. There’s no decoding step, no sampling temperature, and no need to parse a response like “I think the answer is C because…”.
Why this matters: for any workflow where your system must pick from allowed actions—think tool selection, routing, moderation labels, or retrieval choice—you get calibrated scores you can threshold, log, and audit. It’s closer to a classifier’s behavior with the capacity of a large language backbone.
Key takeaway: Treat some tasks as decisions, not generations—and you may gain reliability, speed, and simpler evaluation.
Results at a glance (and what they imply)
On Decision Index 0.3, the v2.1 models improved over v2 on aggregate: the 27B moved from 55.11 to 57.03, and the 9B from 44.09 to 47.20. Gains concentrate in Knowledge, Language, and Retrieval, which tracks with the decision framing: when the options are known, retrieval and understanding can be disentangled from generation quirks.
The notable downside is a Tools regression, with a concentrated drop on When2Call (down 6.38/7.44 for 27B/9B). The author doesn’t claim statistical significance for per-category deltas, so treat all changes as directional until replicated.
For developers, the pattern suggests a few hypotheses worth testing:
- Are the tool-call options or schemas sufficiently represented during training?
- Does the decision framing over-penalize partial credit tool plans compared to end-to-end generation?
- Could calibration tweaks (temperature scaling, isotonic regression) recover some tool-use accuracy without retraining?
Data hygiene and evaluation transparency
The contamination sweep is refreshingly explicit: public benchmark families overlapping public training sources were scanned via text matching with no direct hits. That’s not the end of the story—paraphrases could still leak—but it narrows the risk. The submission used the official scorer and was sent to the Decision Index board, which reportedly adds private tests before ranking, a healthy practice that more evals should emulate.
For those who’ve ever wondered whether a regression came from data drift or metric drift, this setup offers a replicable baseline. You can pull the reported results from the Hugging Face dataset and compare them with in-house reruns.
Quantization notes: how close is close enough?
Quantization is often where theory meets ops. Here, agreement with full weights on 260 held-out questions per build is reported, with examples like the 27B Q8_0 at 259/260 and the 9B MLX 4-bit at 237/260. Per-question records are published, which is exactly the kind of artifact teams need for regression triage.
In practice, closed-set logits are quantization-friendly: you care about rank order and calibrated probabilities across a small label set, not long-form generation quality. If you’re deploying on GPUs with CUDA or on CPU edge boxes, these agreement numbers could translate to simpler infrastructure decisions. The codebase targets common PyTorch stacks, and the weights are hosted on Hugging Face, so getting to a first inference pass should be straightforward.
Why this matters for builders
- Tool routers and function call controllers: Instead of prompting a generator to “choose a tool,” score tools directly and set thresholds. Log the full probability vector for auditability.
- RAG and retrieval choice: Treat “Which chunk to fetch?” as a decision. Use the probabilities to guide top-k strategies and fallback logic.
- Content moderation and policy labeling: Closed-set outputs with calibrated probabilities can map cleanly to enforcement tiers.
- Education and assessment: Multiple-choice questions become native workloads—no answer-string parsing needed.
- Safety-critical gating: When you must prove what the model could have done, a compact vector of option probabilities is easier to reason about than open-ended text.
Compared to a general generator, you trade flexibility for control. For many production systems, that’s the right call.
Open questions and suggested experiments
- Tool-use regression: Is the drop on
When2Calla calibration issue or a representation gap? Try post-hoc calibration (e.g., temperature scaling) and see if the Tools score recovers. - Metric diversity: Add Brier score and Expected Calibration Error (ECE) alongside accuracy. Decision systems live or die by calibration.
- Robustness to paraphrase: Since the contamination check is text-match based, run a paraphrase-augmented audit to estimate upper bounds on leakage.
- Ablations on option set size: How do probabilities and confidence behave as you scale from 3 to, say, 20 labels?
How to try it
The project keeps setup simple. Here’s a quick path:
- Grab the weights on Hugging Face — 27B: jebadiah-27b-v2-1, 9B: jebadiah-9b-v2-1.
- Review code and evals on GitHub. The repository outlines data formatting: structured context, typed question, fixed options.
- Run a minimal inference pass: feed your option list and read the per-label probabilities without any decoding step.
- Reproduce the published eval slice: pull the results dataset and compare with your hardware.
Helpful shortcuts for local experiments:
- git clone the repo and start with a small, fixed option set to verify probability calibration.
- Try a quantized build first to validate deployment constraints; fall back to full precision if a critical label flips.
The bigger picture
Jebadiah v2.1 is part of a broader trend: right-sizing models to the job. Not every task needs a free-form generator. For decision-heavy systems, logit-scored labels can yield cleaner metrics, auditable behavior, and simpler ops. The reported gains on Knowledge/Language/Retrieval, the transparent contamination notes, and the quantization agreement records make this release especially actionable for practitioners.
If your roadmap includes routing, tool calling, or high-stakes multiple choice, this is worth a spin. The models are open (Apache-2.0), the artifacts are published, and the approach invites replication and critique. That’s the kind of engineering posture the community can build on.
“Decisions over generations” isn’t just a style choice—it’s an architecture bet. If your product’s critical path is a label, not a paragraph, this paradigm could be the fit.
Recommended Resources
As an Amazon Associate, I earn from qualifying purchases.