By task
Fact-check an AI answer before you use it
Paste what the chatbot told you. Get a verdict you can trace, sentence by sentence, to a real source.
You asked a chatbot for background on a company, a court case or a news story, and it gave you three tidy paragraphs. They read well. You're about to paste them into a client email, a board note or a slide. Then one sentence snags: a dollar figure, a date, a word like "full" or "first" that you can't place. This page shows how to check an AI answer claim by claim, what CiteJury gives you back, and a real example where a one-line summary of a widely reported story was wrong in the one detail a reader would repeat.
Where AI answers go wrong
Chatbots are good at sounding right. An answer reads like a briefing note whether the facts underneath are solid, stale or invented. You don't catch a fluent paragraph by rereading it. You catch it by checking each claim against a source.
The problem is measurable. In October 2025 the European Broadcasting Union and the BBC published a study of more than 3,000 answers from four popular AI assistants, reviewed by journalists at 22 public service broadcasters in 18 countries. 45% of the answers had at least one significant issue. Sourcing was the biggest single problem, at 31%, and 20% had major accuracy issues, including invented details and outdated information.
The damage doesn't stay in the chat window. In 2025 Deloitte Australia agreed to repay the last instalment of an A$440,000 contract after a report it wrote for the Department of Employment and Workplace Relations was found to contain a fabricated quote from a federal court judgment and references to research papers that don't exist.
It's also a story that's easy to compress. "Deloitte refunded part of its fee" and "Deloitte refunded its fee" are a word apart, and only one is true.
How CiteJury checks an answer
If you need to rely on one sentence, paste it into Check a claim. For the whole answer, choose Check a document and paste the text. CiteJury finds every checkable claim in it, and you pick the ones that matter (up to 25) and can reword them before anything runs.
Each claim is split into parts, so a sentence with a name, a number and a date gets a verdict on each. CiteJury searches several independent search indexes, favors primary sources such as filings, regulators, official pages and wire reports, and looks for evidence against the claim as well as for it. The best passages are locked into a sealed evidence pack with an ID. Nothing can be added to it later.
In a Standard check, Claude, ChatGPT and Grok each read that same pack and answer separately. None of them can search the web at that stage, and every quote an AI cites is checked word for word against its source, so a verdict can't rest on what a model remembers. A final review settles disagreements and writes the verdict: Correct, Incorrect, Sources disagree or Can't confirm.
That last outcome matters. Asking the chatbot that wrote the answer whether it's sure gets you a second answer from the same place as the first. Here, if the evidence doesn't settle a claim, you get Can't confirm, not a guess.
This is what a finished check looks like, on a line that turns up in plenty of AI answers and dinner-table arguments: that humans use only 10% of their brains.
“Humans only use 10% of their brains.”
Incorrect
Claude, ChatGPT and Grok agree12 sources usedTook 86 seconds
It's a myth. Neuroscience sources say the whole brain is used.
Humans use only 10% of their brains
MIT, Johns Hopkins and UW call it a myth.
The other 90% is inactive
Imaging shows activity across the whole brain.
Safe way to say it
Imaging shows the whole brain is active, though not every region is busy at the same moment.
All three AIs said Incorrect in 86 seconds, with MIT's McGovern Institute and Johns Hopkins among the 12 sources. The safe way to say it, at the bottom, is a version of the sentence the evidence does support, ready to paste.
The Deloitte refund, checked
We checked the kind of summary an assistant might give you if you asked what happened:
Deloitte refunded the full A$440,000 fee to the Australian government over its AI-error report.
CiteJury split it into four parts: that Deloitte refunded money over a report with AI-generated errors, that the contract was worth A$440,000, that the refund was the full fee, and that the report contained errors from generative AI. It searched 94 results, with the Guardian, the Australian Financial Review and the ABC among them, and locked 12 sources into the evidence pack. Here's that run, stage by stage:
1Read the claim
Splitting it into parts…
2Find sources
0
results searched
Primary sources first
3Lock the evidence
Nothing can be added later
4Three AIs answer
Same evidence, answered separately
5Verdict
IncorrectAll 3 AIs agree. The refund was partial, the final instalment, not the full A$440,000.
Deloitte agreed to a partial refund, the final instalment of its A$440,000 contract.
Every quote checked word for word
The claim: Deloitte refunded the full A$440,000 fee to the Australian government over its AI-error report.
Three of the four parts held. The story is real, the contract figure is right, and Deloitte acknowledged using AI. The fourth part didn't: the sources say Deloitte repaid only the final instalment of the contract, not the full fee. Claude, ChatGPT and Grok each reached Incorrect on their own, and the check took a little over three minutes.
Every source gets its own ruling: the Guardian's report is marked Contradicts it, the ABC's Backs it up and the AFR's Background. You can open each one to read the passage behind the ruling.
The headline reads: "The story is real, but the refund was partial, not the full fee." And the safe way to say it:
Deloitte agreed to a partial refund, the final instalment of its A$440,000 contract, not a repayment of the full fee.
Later reporting put the repayment at just over A$97,000, well under a quarter of the contract. The gap between "full" and "final instalment" is exactly the kind of detail a fluent summary drops, and exactly the one a reader would quote back to you.
Check an AI answer yourself
- Mark what you'll rely on. Skim the answer and pick out the sentences that carry facts: names, numbers, dates, quotes, and anything that says first, largest, only or full.
- Pick the mode. For one or two claims, paste each into Check a claim. For a whole answer, use Check a document, paste the text and choose up to 25 of the claims it finds.
- Paste the wording as written. Don't tidy it first. If the chatbot said "the full fee", check "the full fee". The wording is what you're about to repeat.
- Choose a tier. Quick uses one AI with no final review, about 2 credits, and suits a first pass. Standard uses Claude, ChatGPT and Grok plus the final review, about 8 credits, and usually takes one to three minutes. Deep searches wider and lets the AIs run their own searches while sources are gathered, about 32 credits, which helps with niche or recent topics.
- Read the parts, then the sources. The headline tells you what's wrong. The parts tell you which piece. Open anything ruled Contradicts it.
- Use the safe wording. Copy it, or edit your sentence to match what the sources support. If something's unclear, ask a follow-up question. It answers only from the same sources and costs about 1 to 2 credits.
What to check first
You rarely need to check every sentence. Start where answers slip most often and where your reader is most likely to act.
- Qualifiers that shrank or grew. Partial becomes full, estimated becomes reported, up to becomes about. The Deloitte example turned on one word like this.
- Numbers without a year. A figure that was right in 2023 can be wrong now. See how to trace a statistic back to its source.
- Quotes and attributions. Anything in quotation marks attributed to a real person or a court deserves a check. An invented quote from a judgment was part of what went wrong in the Deloitte report.
- Citations the chatbot supplies. A link or a paper title isn't evidence until someone has opened it and found the claim there.
- Cause and effect. Summaries sometimes flip the order of events. Check which came first.
If the AI text is going into a longer report or deck, it's usually quicker to check the whole document once it's drafted. Students and researchers using an assistant for drafts can read about checking the facts in a paper before you submit.
What it won't do
- It can be wrong. A verdict is only as good as the sources the search finds, and results can be wrong or incomplete. That's why every sentence links to its source. Check the ones you'll rely on.
- It can't read every page. Pages behind paywalls or logins, and some sites that block automated access, can't be read. CiteJury lists the pages it couldn't open so you know where to look yourself.
- It won't tell you whether text was written by AI. It checks whether what a sentence claims holds up against sources. It isn't an AI-text detector or a plagiarism checker.
- It checks text only. Images, charts and video in an answer aren't checked.
- Very recent events are hard. If something happened this week, coverage may be thin, and Can't confirm is the likely result.
- It isn't advice. A verdict on a legal, financial or medical claim isn't legal, financial or medical advice.
AI answers make a useful first draft. The trouble starts when a first draft goes out as a final one. Checking the claims you'll rely on takes minutes and leaves you with a source for each and wording you can defend.
Questions people ask
Can I paste a whole ChatGPT answer, or only one sentence?
Both. Paste a single sentence into Check a claim, or paste the whole answer into Check a document, which finds every checkable claim and lets you pick up to 25 to check.
Does CiteJury use ChatGPT to check ChatGPT?
In a Standard check, Claude, ChatGPT and Grok each read the same sealed pack of sources and answer separately, with no web search at that stage, and a final review settles any disagreement. The verdict rests on quotes from sources you can open, not on what a model remembers, and you can turn any of the three AIs off in Settings.
What happens if the sources don't settle a claim?
You get Can't confirm instead of a guess. The result also lists the pages CiteJury couldn't open, so you know where to look yourself.
Does it work on answers from Gemini, Claude or Perplexity?
Yes. CiteJury checks the text you paste, whichever chatbot wrote it.
How much does it cost to check an AI answer?
Credits cost $0.10 each, and you pay only what a check actually uses: an estimate is held while it runs and the rest comes back. Failed checks are free, and new accounts get free credits to start.
Check the answer before you send it.
New accounts get 100 free credits. Failed checks cost nothing.