Try “car wash”, “subscription box”, “Austin” · Esc to close

JevBench

Reproducible benchmark for typed decision models versus LLMs.

Developer tool / API SaaS & software Show HN · launch post · ▲ 153

Visit site

benchmarkheaven.com

What it does

JevBench is a benchmark for evaluating Jev-class decision models—systems that return bounded choices and probabilities rather than open-ended text. The benchmark runs 1,624 decisions per system (904 open and 720 sealed) and measures three dimensions: intelligence (chance-corrected accuracy), calibration (confidence quality), and a composite capability score that can weight accuracy, latency, and cost according to user preferences.

The tool ranks systems side by side, allowing direct comparison of how models perform when cost and latency are constrained. Users can set cost and latency caps relative to a reference model and see how different systems perform within those constraints.

Who it is for

JevBench targets founders and engineers evaluating decision models as alternatives to large language models. The benchmark is designed for buyers comparing Jev-class systems on grounds of speed, cost, and accuracy together. It is also relevant to anyone building or integrating systems where bounded outputs (not text generation) are the goal.

Pricing

The site does not show prices.

How it stands out

JevBench positions itself as a reproducible comparison tool specifically for Jev-class models rather than general-purpose LLMs. The benchmark uses a weighted scoring system that lets users configure how much they care about intelligence, calibration, latency, and cost. The launch text claims Jev-class models are "disruptively faster and cheaper than LLMs" while matching them on accuracy for text input.

The benchmark is transparent about its methodology: it publishes 1,624 test decisions, keeps half of them sealed to prevent overfitting, and provides aggregate results with SHA256 checksums. Users can adjust cost and latency caps to see how rankings shift under different constraints.

What a founder should check

  1. Incumbent alternatives. Verify what systems founders currently use to make bounded decisions. Are they comparing against LLMs, traditional decision trees, or other typed-output systems? What switching costs exist?
  1. Benchmark gaming risk. Examine whether the sealed decision set (half the test cases) is large enough to prevent model optimization to the public 904 decisions. Ask whether competitors disclose their own benchmark runs versus being ranked without consent.
  1. Market demand. Confirm that buyers actually optimize for the three weighted dimensions (accuracy, speed, cost). Check whether most decision-making workloads fit the "state plus bounded rubric" constraint, or whether open-ended reasoning is still the primary need in target use cases.

Thinking of building something like this?

Every launch here is a competitor to somebody's idea. If yours is close, check it against the market before you build: the Full Check names the rivals, the prices and the gaps.

Check an idea like this

More developer tool / api launches

All

Wispbit

Linter that enforces codebase standards with AI coding agents.

Developer tool / API SaaS & softwareShow HN ▲ 31

OnlyJPG

Private browser-based converter for any image format to JPG.

Developer tool / API SaaS & softwareShow HN ▲ 64

Duck-UI

Browser-based SQL IDE for DuckDB running entirely in WebAssembly.

Developer tool / API SaaS & softwareShow HN ▲ 213

Checked ideas in SaaS & software