How to Measure a Legal AI Tool's Citation Accuracy Yourself (India, 2026)
A reproducible method for testing any legal AI tool’s citation accuracy yourself, in an afternoon, with a query set, a four-point check, and a written rubric.
How To · Test Legal AI Accuracy
You do not need a lab, or a vendor’s word, to find out how accurate a legal AI tool actually is. You can test it yourself, in an afternoon: build a small set of real research questions, run them through the tool, check each citation it returns against a primary source by hand, and score the result against a rubric you wrote down before you started. This page sets out that method step by step, so you can run it on any tool, including Claw, and trust the number because you produced it, not because someone sold it to you. Last verified 17 September 2026.
- You can test it yourself: build 20 to 50 real research queries, run them, hand-check every citation, and score it against a written rubric.
- The four-point check: does the case exist, is the court and date right, does it support the proposition claimed, and is it still good law.
- A wrong proposition is a failure, even if the case itself is real and correctly cited.
- Distrust a claim with no denominator, no shared query set, or no independent rubric.
- Worked example: the Claw citation accuracy benchmark follows this exact method, with the query set being prepared for public download.
01Why you should test this yourself
Almost every legal AI tool sold in India, and internationally, claims to be accurate. Very few show you a number, a sample size, and a date you can check that claim against. That gap is not a small thing for a working lawyer, because the person who eventually pays for a bad citation is not the vendor, it is the advocate who filed it.
A vendor’s claim is marketing until you can check it
A feature list tells you what a tool is supposed to do. It does not tell you how often it gets a citation right. Coverage claims, court counts, and "AI-powered" language are all useful context, but none of them are a substitute for actually running the tool on real questions and checking what comes back.
The failure mode is well documented
This is the same problem behind AI-generated citations that do not exist, a risk that is now well documented rather than theoretical (see our explainer on whether legal AI hallucinates citations in India). A general-purpose AI model with no grounding in real judgment text can produce a citation that reads perfectly and refers to nothing real. The only way to know whether a specific tool avoids this is to test it against real questions from your own practice.
Testing yourself also fits how you actually work
A vendor’s own benchmark, even an honest one, is built on their choice of questions. Your test set is built on your practice areas, the courts you actually appear in, and the kind of research question you ask every week. That makes the result directly useful to your own decision, in a way a generic benchmark cannot be.
The method in one line
Build 20 to 50 real research questions from your own practice, run them through the tool, hand-check every returned citation against a primary source using a four-point test, score it against a rubric you wrote before you started, and report the percentage together with your sample size and the date you ran it.
02The method, in five steps
This is the same method behind any credible, checkable accuracy figure. None of these steps need special tools, only time and discipline.
- Build a query set of real research questions. Take somewhere between 20 and 50 questions from matters you have actually worked on, spread across your practice areas and the courts you actually deal with. Avoid artificially easy lookups (a famous, heavily cited case) and avoid trick questions designed to fail the tool. The goal is a fair, realistic sample, not a stress test in either direction.
- Run every query through the tool, the same way each time. Use the same wording and the same settings for each question. If you are comparing more than one tool, run the identical query set through each one, so the comparison is fair. If you want to test at volume rather than typing each query by hand, some tools also offer programmatic access. See litigation monitoring APIs if your evaluation needs to run at scale rather than query by query.
- Apply the four-point check to every citation returned. For each citation the tool gives you, check it against a primary source: does the case exist, is the court and date right, does the judgment actually support the proposition, and is it still good law. The full check is set out in the next section.
- Score each result against a rubric you wrote down first. Decide, before you see the results, exactly what counts as a pass and what counts as a fail. A citation that exists but is quoted for the wrong proposition is a failure, not a partial success. Writing the rubric down first stops you from quietly moving the goalposts once you see how the tool actually performed.
- Report the percentage with your sample size and the date. "94 percent accurate" means nothing on its own. "94 percent, out of 35 queries, tested on 12 September 2026" is a number someone else can question, repeat, and check. Keep the list of failures too. A result with no failures in it is usually a sign the test was too easy, not that the tool is perfect.
A percentage with no stated sample size is not a measurement. It is a marketing line wearing a number as a costume.
03The four-point citation check
This is the same check used to verify every citation on this site’s own benchmark page, and it works on any tool’s output, not just Claw’s.
| What you check | The question you are answering | Primary source to check against | Typical time |
|---|---|---|---|
| Citation exists | Does this case exist, in this court, with this citation, in this year? | Indian Kanoon or the court’s own website | Under 1 minute |
| Court and date correct | Are the bench, date, and citation number exactly right? | The court’s own judgment copy or cause list | Under 1 minute |
| Proposition matches | Does the judgment actually say what the AI claims it says? | The full judgment text | 1 to 2 minutes |
| Still good law | Has it been overruled, distinguished, or stayed since? | SCC Online, Manupatra, or Indian Kanoon citing references | 1 to 2 minutes |
A citation only passes if it clears all four. A case that genuinely exists, with the right court and date, but is cited for a proposition it does not actually support, still fails the check. This is the point where most self-graded accuracy claims quietly go soft, because "the case is real" and "the citation is correct" are not the same thing.
04How to score it honestly
The rubric is what turns a spot check into a real measurement. Write it down before you run a single query, and do not change it once you have seen the results.
- Pass: the citation exists, the court and date are correct, the proposition it is cited for is actually supported by the judgment, and it is still good law as far as you can confirm.
- Fail: any of the above is wrong, including a real case cited for the wrong proposition. Do not create a middle "partially correct" category. It invites the result to drift toward whatever number looks best.
- Unclear, not a pass: if you genuinely cannot confirm a point (for instance, whether a judgment has since been distinguished) within a reasonable time, mark it as unresolved and say so in your report, rather than counting it as a pass by default.
- Whoever scores it, use the same rubric on every tool you test, including the one you currently use or are inclined to buy. A rubric that is easier on your preferred tool is not a measurement, it is a preference dressed up as one.
When you report your result, publish the failures alongside the percentage, even briefly. A benchmark with no losses in it usually means the test was too easy, the rubric was too soft, or both.
05Distrust triggers in a vendor’s own benchmark
When any vendor, Claw included, publishes an accuracy claim, three things should make you pause before you believe it.
- No denominator. A bare "99%+ accurate" with no stated sample size tells you nothing. Ninety-nine correct out of a hundred is a real result. Ninety-nine correct out of ninety-nine, hand-picked, is not the same claim at all, and there is no way to tell which one you are looking at without the number of queries tested.
- The query set is described but never shared. If a vendor talks about "a realistic set of legal research questions" but will not let you see or run that set yourself, you cannot check the claim, only trust it. A query set you can download and re-run is what turns a claim into a checkable measurement.
- Only the vendor grades it, with no stated rubric. If nobody outside the vendor scored the result, and there is no written definition of what counts as a pass or a fail, the number reflects the vendor’s judgment about their own product, not an independent check.
None of this means a vendor-published number is automatically wrong. It means you should apply the same four-point check and the same rubric from this page to a sample of their claimed results yourself, before you rely on the headline figure for a buying decision.
This applies to Claw too
Any accuracy figure Claw publishes should meet the same test: a stated sample size, a query set you can check or reproduce, and a rubric written down in advance. If you run this method against Claw and get a different number, that is useful information, not something to hide.
06A worked example you can copy
If you want to see this method applied end to end, rather than described in the abstract, our citation accuracy benchmark for Claw is a worked example that follows exactly the method on this page: a defined query set, the same four-point check, a rubric fixed before scoring, a stated sample size and date, and a disclosed list of failure categories rather than a claim of perfection.
That page also notes that the underlying query set is being prepared for public download, specifically so any firm can re-run the identical test, against any tool, and check the published number for themselves rather than take it on trust. That is the standard this page is asking you to hold every legal AI accuracy claim to, including ours.
07How many queries do you actually need
Somewhere between 20 and 50 real queries is usually enough for a first, useful test. Fewer than that, and a couple of unlucky results can swing your percentage by several points, which makes the number unstable. More than that, and you are usually better off spending the extra time widening the practice areas covered rather than repeating similar questions.
What matters more than the exact count is spread: cover the courts you actually appear in, the practice areas you actually work, and a mix of older and recent judgments, since recency and obscure party names are where most tools struggle most. A tool that scores well on twenty well-chosen, realistic questions across your real practice tells you more than a tool that scores well on a hundred easy ones.
If your evaluation needs to run continuously, or against a large volume of matters rather than a one-time check, that starts to become a different job, closer to ongoing litigation monitoring than a one-time accuracy test. For a one-off evaluation before you buy or renew a tool, the method above is enough.
08Where Claw fits
Claw is an all-in-one legaltech platform for Indian advocates, law firms, and corporate legal teams, combining AI-based case search, an AI legal assistant (Legal GPT), case management, and compliance automation across all Indian courts and tribunals.
Claw’s position on this is simple: the method on this page is the same one used to produce its own published citation accuracy figure, and the same one any firm evaluating Claw, or an alternative such as the other AI legal research tools available in India, should run for themselves rather than accept on trust. Because Claw’s case search is grounded in real judgment text, with the source shown alongside the answer, its citations are built to survive exactly the four-point check described above. The honest position is not "trust our number", it is "here is the method, run it yourself, and if you get a different result, tell us."
09Sources and further reading
Primary sources to check citations against, and the tools referenced on this page:
- Indian Kanoon (free public judgment database): indiankanoon.org
- Supreme Court of India (official judgments): sci.gov.in
- SCC Online: scconline.com
- Manupatra: manupatra.com
- Claw: clawlaw.in
This method works for any legal AI or legal research tool, not only the ones named on this page.
10Frequently asked questions
How do I know if a legal AI tool is accurate?
Test it yourself rather than relying on a vendor’s claim. Build a set of real research questions from your own practice, run them through the tool, check every returned citation against a primary source like Indian Kanoon or the court’s own copy, and score the result against a rubric you wrote before you started. A checkable number you produced yourself is worth more than any marketing figure.
How many queries do I need to test a legal AI tool?
Somewhere between 20 and 50 real queries is usually enough for a first, useful test. Fewer than that and a couple of unlucky results can swing your percentage significantly. What matters more than the exact count is covering the courts and practice areas you actually work in, rather than repeating similar or artificially easy questions.
What counts as a wrong citation?
A citation is wrong if it fails any part of the four-point check: the case does not exist as cited, the court or date is wrong, the judgment does not actually support the proposition it is cited for, or it is no longer good law. A real case cited for the wrong proposition should be scored as a failure, not a partial success, even though the case itself exists.
Can I trust a vendor’s own accuracy claim?
Treat it as a starting point, not a conclusion, until you can check it. Watch for three warning signs: no stated sample size, a query set that is described but never shared so you can reproduce it, and grading done only by the vendor with no written rubric. A vendor claim that avoids all three is worth more attention than a bare percentage with none of that detail.
Should I test the tool I already use, or only the ones I am considering buying?
Test all of them, including your current tool, using the identical query set and rubric. This is the only way to know whether switching, or staying, is actually the better decision, rather than assuming your current tool is fine because you have not checked it.
Where can I see this method applied to a real result?
The Claw citation accuracy benchmark is a worked example that follows this exact method: a defined query set, the same four-point check, a rubric fixed in advance, a stated sample size and date, and a disclosed list of failures rather than a claim of perfection. The underlying query set is being prepared for public download so the test can be reproduced by anyone.