Skip to main content

Definition

The Jev-as-a-judge test scores every row against criteria you write in plain language. In development that is the validation set. In monitoring it is the production window. A calibrated classifier returns a verdict for each row, so the test can cover the whole set instead of a sample. LLM-as-a-judge sends a sample of rows to a generative model and asks it to grade them and explain the grade. The project sample is often about 50 rows. Jev-as-a-judge judges every row unless you opt into a sample, and it returns a score rather than a written explanation.

Taxonomy

  • Task types: LLM.
  • Availability: and .

Why it matters

  • A sample can miss a failure that shows up on only a few rows. Judging every row means those failures count.
  • Several criteria share one request per row, so adding a criterion costs little beyond the first.
  • You choose whether the test watches the share of rows that pass, or the average verdict. On a range criterion those two numbers tell different stories.

When to use it

Use Jev-as-a-judge when you want a verdict on every row and a generative judge would be too expensive to run that way. Use LLM-as-a-judge when you need a written rationale, or when the judge has to be a specific LLM provider and model you already connect.

Criteria

A test takes one to five criteria. Each one needs criteria: what the judge should decide about the row, up to 8192 characters. You can also set:
  • name — the display name, up to 50 characters.
  • key — a stable id, up to 50 characters. Scores stay attached to this id, so you can rename the criterion without dropping verdicts already computed for it. Keys on one test must be unique. Set one when you expect to rename or reorder criteria. The measurement name (criteria0PassRate, criteria1MeanScore, and so on) follows list order, not key.
  • kind — noul (the default) or score.
  • threshold — the value a row must reach to count as passing. The default is 0.5. For a binary criterion this is the probability of a 1. For a range criterion it is a position on that criterion’s scale.
  • passesAbove — which side of the threshold passes. The default is true. Set it to false when a higher verdict is worse, such as a severity rubric that should pass below its threshold.
Binary (noul) scores each row 0 or 1. Write criteria so the judge knows what each level means, in the same rubric style as LLM-as-a-judge. Name both scores in the text, for example 1 if the answer is fully supported by the retrieved context, 0 if it is not, or 1 if profanity is flagged, 0 if no profanity. The create form placeholder uses the same shape: describe what the judge should decide, for example 1 if the condition is met, 0 if not. After judging, the row is 1 when the probability of a 1 is on the passing side of threshold, and 0 otherwise. With the defaults, a probability of 0.5 or higher scores 1. Range (score) rates the row on a numeric scale and returns that score. Set rangeMin and rangeMax as inclusive integers. The scale must cover 2 to 10 whole numbers. The default is 0 to 1. The numerals do not mean anything on their own, so say what the ends of the scale mean in criteria. A range verdict can land between the ends. A binary verdict is only 0 or 1. The create form writes a single criterion. Binary is tagged 0, 1 and Range is tagged 0–1. Wider scales, a custom row threshold, and additional criteria are set in tests.json.

Pass rate and mean

Each criterion publishes two measurements. N is its position in the criteria list, starting at 0 (criteria0 … criteria4):
  • criteriaNPassRate — Pass rate, the fraction of judged rows whose verdict cleared that criterion’s threshold. In the create form this threshold, and the mean, accept a value from 0 to 1.
  • criteriaNMeanScore — the average of the raw verdicts.
On a binary criterion every verdict is 0 or 1, so the mean and the pass rate are the same number. On a range criterion they are not. Ninety-five rows at 1.0 and five rows at 0.0 have the same mean (0.95) as one hundred rows at 0.95. If the row threshold is 0.5, the first window has five failures and the second has none. The pass rate shows that. The mean does not. Use the pass rate when any failed row should count, even if the rest of the window scores highly. Use the mean when you care how far the typical verdict sits on the scale. For a pass rate, the test threshold value is a fraction between 0 and 1. For a mean, value is a number on that criterion’s own scale. Rows the judge declines are left out of both measurements. They are not counted as failures.

API key

Jev-as-a-judge reads TYPESAFE_API_KEY from environment variables. Openlayer resolves a key in this order: the project, then the workspace, then a platform key. A platform key still runs the test. Those rows are judged through Openlayer’s TypeSafe account. Set your own key when the data and the spend need to stay on your account, including data-residency and bring-your-own-key requirements. If no key resolves, the test is skipped rather than run. The create form shows which case you are in. With no key, you can still create the test, and it is skipped on every run until you add TYPESAFE_API_KEY. With only the platform key, the form tells you the rows are judged through Openlayer’s TypeSafe account.

Judge and sampling

classifier_evaluator chooses the judge, and how much of the data it sees. Omit it to judge every row with the default model.
  • provider — typesafe. This is the only provider.
  • model — jev-latest (the default) or jev-preview. jev-latest tracks the current stable release. jev-preview is the next release, ahead of general availability. Any other id is rejected, including a partial version such as jev-1.13.
  • sample — omit it to judge every row. Set it only to trade coverage for cost.
    • fraction — the share of rows to judge, from 0 to 1. The default is 1. The same rows stay in the sample if a run is retried or a monitoring window resumes.
    • limit — a cap on rows judged for this test, counted across resumed runs rather than once per run. Omit it for no cap.
In the create form, open Judge settings for Judge model, Sample size (%) (100 judges every row), and Data limit (rows) (empty means no cap).

Test configuration examples

If you are writing a tests.json, here are valid configurations for the Jev-as-a-judge test. The development example judges every row. The monitoring example judges a quarter of the window, capped at 10,000 rows.