How often does an AI model cite a case number we cannot confirm?
We asked three AI models which case law applies. Each question ran twice: once with no help, once with the relevant decisions from our collection. We then looked up every case number they cited in that collection.
45%
168 of 374 cited case numbers not confirmed
0.2%
1 of 471 cited case numbers not confirmed
- Exception: Claude Opus 4.8, Tax law, 1 of 26 not confirmed.
What this means for your work: check every case number an AI tool gives you before you cite it. This page shows how often that is necessary.
Models tested, measured on 21 July 2026: Claude Opus 4.8, GPT-5.6 Sol, Kimi K3
Looked up in our collection of court decisions, 510,000+ today.
As of: 21 July 2026
One question, two runs, verbatim
Aggravated arson, section 306b(2) of the German Criminal Code.
without our data
1 of 4 not confirmed
Case numbers cited, looked up in our collection:
- 4 StR 19/00 confirmed
- 3 StR 541/14 confirmed
- 5 StR 280/18 confirmed
- 2 StR 213/20 not confirmed
Nach ständiger BGH-Rechtsprechung reicht eine bloß abstrakte Lebensgefahr – etwa allein der Umstand, dass sich Menschen in dem angezündeten Gebäude befinden – nicht aus.
with our data
0 of 7 not confirmed
Case numbers cited, looked up in the decisions we handed over:
- 4 StR 432/18 confirmed
- 1 StR 416/17 confirmed
- 3 StR 172/17 confirmed
- 3 StR 293/12 confirmed
- 3 StR 45/13 confirmed
- 4 StR 399/17 confirmed
- 3 StR 336/13 confirmed
Gefährdungsvorsatz liegt vor, wenn der Täter die Umstände erkennt, aus denen sich die konkrete Todesgefahr ergibt, und sich mit dem Eintritt dieser Gefahr abfindet.
From the same measurement run as every figure on this page. Model: GPT-5.6 Sol. Measured on 21 July 2026.
BGH stands for Bundesgerichtshof, the highest German court in civil and criminal matters. Case numbers such as 4 StR 19/00 are the docket numbers German courts assign to a case. StR marks a criminal division of the BGH. The two answers above are in German, as the model wrote them.
Results by area of law
- Tax law73%48 of 66without our data not confirmedwith our data:1 of 75with our data not confirmed
- Social security law52%23 of 44without our data not confirmedwith our data:0 of 63with our data not confirmed
- Employment law50%38 of 76without our data not confirmedwith our data:0 of 77with our data not confirmed
- Criminal law43%21 of 49without our data not confirmedwith our data:0 of 80with our data not confirmed
- Business crime28%16 of 57without our data not confirmedwith our data:0 of 67with our data not confirmed
- Tenancy law27%22 of 82without our data not confirmedwith our data:0 of 109with our data not confirmed
Results by model
Where each model goes wrong
Without our data, how often a model goes wrong depends on the area of law. Each square is one case number the model cited. Red with a cross: not confirmed. The page for each area shows the single questions.
Without our datanot confirmed, checked against our whole collection
Kimi K3: 17 of 23 not confirmed
Claude Opus 4.8: 19 of 25 not confirmed
GPT-5.6 Sol: 12 of 18 not confirmed
With our datanot confirmed, checked against the decisions provided
Kimi K3: 0 of 27 not confirmed
Claude Opus 4.8: 1 of 26 not confirmed
GPT-5.6 Sol: 0 of 22 not confirmed
Without our datanot confirmed, checked against our whole collection
Kimi K3: 11 of 14 not confirmed
Claude Opus 4.8: 5 of 14 not confirmed
GPT-5.6 Sol: 7 of 16 not confirmed
With our datanot confirmed, checked against the decisions provided
Kimi K3: 0 of 24 not confirmed
Claude Opus 4.8: 0 of 24 not confirmed
GPT-5.6 Sol: 0 of 15 not confirmed
Without our datanot confirmed, checked against our whole collection
Kimi K3: 19 of 28 not confirmed
Claude Opus 4.8: 10 of 22 not confirmed
GPT-5.6 Sol: 9 of 26 not confirmed
With our datanot confirmed, checked against the decisions provided
Kimi K3: 0 of 32 not confirmed
Claude Opus 4.8: 0 of 30 not confirmed
GPT-5.6 Sol: 0 of 15 not confirmed
Without our datanot confirmed, checked against our whole collection
Kimi K3: 15 of 20 not confirmed
Claude Opus 4.8: 2 of 15 not confirmed
GPT-5.6 Sol: 4 of 14 not confirmed
With our datanot confirmed, checked against the decisions provided
Kimi K3: 0 of 32 not confirmed
Claude Opus 4.8: 0 of 25 not confirmed
GPT-5.6 Sol: 0 of 23 not confirmed
Without our datanot confirmed, checked against our whole collection
Kimi K3: 14 of 18 not confirmed
Claude Opus 4.8: 1 of 19 not confirmed
GPT-5.6 Sol: 1 of 20 not confirmed
With our datanot confirmed, checked against the decisions provided
Kimi K3: 0 of 22 not confirmed
Claude Opus 4.8: 0 of 25 not confirmed
GPT-5.6 Sol: 0 of 20 not confirmed
Without our datanot confirmed, checked against our whole collection
Kimi K3: 21 of 32 not confirmed
Claude Opus 4.8: 1 of 21 not confirmed
GPT-5.6 Sol: 0 of 29 not confirmed
With our datanot confirmed, checked against the decisions provided
Kimi K3: 0 of 39 not confirmed
Claude Opus 4.8: 0 of 37 not confirmed
GPT-5.6 Sol: 0 of 33 not confirmed
How we measured
What does “without our data” mean?
- Question
- AI model
- Answer with case numbers
The model answers the question on its own. No search, no tools, no internet.
Checked againstOur whole collection of court decisions, 510,000+ today
What does “with our data” mean?
- Question and relevant decisions
- AI model
- Answer with case numbers
The same model receives a selection of relevant decisions from our collection for the same question. It does not search for them itself.
Checked againstOnly the decisions the model was given
- What counts as one?
- One mention of one case number. If a model cites the same decision twice, we count two.
- How large is the sample?
- Six areas of law, three questions each, three models: 374 case numbers cited without our data and 471 with it. With our data the models cite more case numbers, because they quote from the decisions they were given.
- How often did we measure each question?
- Once per model and run. We take no average and claim no repeat runs.
- Who wrote the questions, and are they public?
- We wrote them ourselves. We do not publish them, so that nobody can tune a model to these exact wordings. We show one question in full below.
What this benchmark does not show
Not confirmed does not mean invented
“Not confirmed” means we did not find the case number in our collection. It does not mean the decision does not exist. It can be genuine and missing from our holdings.
We looked up 16 of these case numbers, from two of the six areas, on dejure.org. Five decisions exist there and are missing from our holdings. Eleven we did not find there either.
- It does not test our search. It tests whether a model cites real references once it is handed the relevant decisions. We assembled that selection by hand for each question.
- The figures apply to the model versions tested on the day of the measurement. They say nothing about later versions.
- All figures come from one measurement written down on 21 July 2026. They do not change as our collection grows. A new measurement gets a new date.
References that hold up, in your work
Klaracase names the decision behind every answer. In research, and through an interface for your own applications.
Talk to us
Law firms and companies
Want to know whether Klaracase fits your work? Write to us. We will then set a time for a call.
Email: info@klaracase.de
Research and press
Questions about the method or the data behind this benchmark? Write to us. We share the measurement data with researchers on request.
Email: pr@klaracase.de