Benchmark

How often does an AI model cite a case number we cannot confirm?

We asked three AI models which case law applies. Each question ran twice: once with no help, once with the relevant decisions from our collection. We then looked up every case number they cited in that collection.

without our data

45%

168 of 374 cited case numbers not confirmed

with our data

0.2%

1 of 471 cited case numbers not confirmed

  • Exception: Claude Opus 4.8, Tax law, 1 of 26 not confirmed.

What this means for your work: check every case number an AI tool gives you before you cite it. This page shows how often that is necessary.

Models tested, measured on 21 July 2026: Claude Opus 4.8, GPT-5.6 Sol, Kimi K3

Looked up in our collection of court decisions, 510,000+ today.

As of: 21 July 2026

One example

One question, two runs, verbatim

This is one of the eighteen questions, shown in full. On the left the answer without our data, on the right with it. The case numbers appear exactly as the model wrote them.
§ 306b Abs. 2 StGBQuestion 1, criminal law, topic arson

Aggravated arson, section 306b(2) of the German Criminal Code.

without our data

1 of 4 not confirmed

Case numbers cited, looked up in our collection:

  • 4 StR 19/00 confirmed
  • 3 StR 541/14 confirmed
  • 5 StR 280/18 confirmed
  • 2 StR 213/20 not confirmed
Nach ständiger BGH-Rechtsprechung reicht eine bloß abstrakte Lebensgefahr – etwa allein der Umstand, dass sich Menschen in dem angezündeten Gebäude befinden – nicht aus.

with our data

0 of 7 not confirmed

Case numbers cited, looked up in the decisions we handed over:

  • 4 StR 432/18 confirmed
  • 1 StR 416/17 confirmed
  • 3 StR 172/17 confirmed
  • 3 StR 293/12 confirmed
  • 3 StR 45/13 confirmed
  • 4 StR 399/17 confirmed
  • 3 StR 336/13 confirmed
Gefährdungsvorsatz liegt vor, wenn der Täter die Umstände erkennt, aus denen sich die konkrete Todesgefahr ergibt, und sich mit dem Eintritt dieser Gefahr abfindet.

From the same measurement run as every figure on this page. Model: GPT-5.6 Sol. Measured on 21 July 2026.

BGH stands for Bundesgerichtshof, the highest German court in civil and criminal matters. Case numbers such as 4 StR 19/00 are the docket numbers German courts assign to a case. StR marks a criminal division of the BGH. The two answers above are in German, as the model wrote them.

Per model

Results by model

All six areas of law combined, per model. The names are the exact model versions tested, measured on 21 July 2026.
  1. Kimi K372%97 of 135without our data not confirmedwith our data:0 of 176with our data not confirmed
  2. Claude Opus 4.833%38 of 116without our data not confirmedwith our data:1 of 167with our data not confirmed
  3. GPT-5.6 Sol27%33 of 123without our data not confirmedwith our data:0 of 128with our data not confirmed

Where each model goes wrong

Without our data, how often a model goes wrong depends on the area of law. Each square is one case number the model cited. Red with a cross: not confirmed. The page for each area shows the single questions.

not confirmedconfirmedEach question ran twice, as two separate answers. That is why a model cites a different number of case numbers.
  1. Tax law

    Without our datanot confirmed, checked against our whole collection

    Kimi K3: 17 of 23 not confirmed

    Claude Opus 4.8: 19 of 25 not confirmed

    GPT-5.6 Sol: 12 of 18 not confirmed

    With our datanot confirmed, checked against the decisions provided

    Kimi K3: 0 of 27 not confirmed

    Claude Opus 4.8: 1 of 26 not confirmed

    GPT-5.6 Sol: 0 of 22 not confirmed

  2. Social security law

    Without our datanot confirmed, checked against our whole collection

    Kimi K3: 11 of 14 not confirmed

    Claude Opus 4.8: 5 of 14 not confirmed

    GPT-5.6 Sol: 7 of 16 not confirmed

    With our datanot confirmed, checked against the decisions provided

    Kimi K3: 0 of 24 not confirmed

    Claude Opus 4.8: 0 of 24 not confirmed

    GPT-5.6 Sol: 0 of 15 not confirmed

  3. Employment law

    Without our datanot confirmed, checked against our whole collection

    Kimi K3: 19 of 28 not confirmed

    Claude Opus 4.8: 10 of 22 not confirmed

    GPT-5.6 Sol: 9 of 26 not confirmed

    With our datanot confirmed, checked against the decisions provided

    Kimi K3: 0 of 32 not confirmed

    Claude Opus 4.8: 0 of 30 not confirmed

    GPT-5.6 Sol: 0 of 15 not confirmed

  4. Criminal law

    Without our datanot confirmed, checked against our whole collection

    Kimi K3: 15 of 20 not confirmed

    Claude Opus 4.8: 2 of 15 not confirmed

    GPT-5.6 Sol: 4 of 14 not confirmed

    With our datanot confirmed, checked against the decisions provided

    Kimi K3: 0 of 32 not confirmed

    Claude Opus 4.8: 0 of 25 not confirmed

    GPT-5.6 Sol: 0 of 23 not confirmed

  5. Business crime

    Without our datanot confirmed, checked against our whole collection

    Kimi K3: 14 of 18 not confirmed

    Claude Opus 4.8: 1 of 19 not confirmed

    GPT-5.6 Sol: 1 of 20 not confirmed

    With our datanot confirmed, checked against the decisions provided

    Kimi K3: 0 of 22 not confirmed

    Claude Opus 4.8: 0 of 25 not confirmed

    GPT-5.6 Sol: 0 of 20 not confirmed

  6. Tenancy law

    Without our datanot confirmed, checked against our whole collection

    Kimi K3: 21 of 32 not confirmed

    Claude Opus 4.8: 1 of 21 not confirmed

    GPT-5.6 Sol: 0 of 29 not confirmed

    With our datanot confirmed, checked against the decisions provided

    Kimi K3: 0 of 39 not confirmed

    Claude Opus 4.8: 0 of 37 not confirmed

    GPT-5.6 Sol: 0 of 33 not confirmed

Method

How we measured

What does “without our data” mean?

  1. Question
  2. AI model
  3. Answer with case numbers

The model answers the question on its own. No search, no tools, no internet.

Checked againstOur whole collection of court decisions, 510,000+ today

What does “with our data” mean?

  1. Question and relevant decisions
  2. AI model
  3. Answer with case numbers

The same model receives a selection of relevant decisions from our collection for the same question. It does not search for them itself.

Checked againstOnly the decisions the model was given

What counts as one?
One mention of one case number. If a model cites the same decision twice, we count two.
How large is the sample?
Six areas of law, three questions each, three models: 374 case numbers cited without our data and 471 with it. With our data the models cite more case numbers, because they quote from the decisions they were given.
How often did we measure each question?
Once per model and run. We take no average and claim no repeat runs.
Who wrote the questions, and are they public?
We wrote them ourselves. We do not publish them, so that nobody can tune a model to these exact wordings. We show one question in full below.
Limits

What this benchmark does not show

Not confirmed does not mean invented

“Not confirmed” means we did not find the case number in our collection. It does not mean the decision does not exist. It can be genuine and missing from our holdings.

11 not found on dejure.org either5 exist there and are missing from our collection

We looked up 16 of these case numbers, from two of the six areas, on dejure.org. Five decisions exist there and are missing from our holdings. Eleven we did not find there either.

  • It does not test our search. It tests whether a model cites real references once it is handed the relevant decisions. We assembled that selection by hand for each question.
  • The figures apply to the model versions tested on the day of the measurement. They say nothing about later versions.
  • All figures come from one measurement written down on 21 July 2026. They do not change as our collection grows. A new measurement gets a new date.

References that hold up, in your work

Klaracase names the decision behind every answer. In research, and through an interface for your own applications.

Contact

Talk to us

Law firms and companies

Want to know whether Klaracase fits your work? Write to us. We will then set a time for a call.

Email: info@klaracase.de

Research and press

Questions about the method or the data behind this benchmark? Write to us. We share the measurement data with researchers on request.

Email: pr@klaracase.de