AI Report Writing: Good at Prose, Bad at Facts
AI report writing produces fluent prose instantly, and fluency was never the hard part. A New York court fined two lawyers $5,000 in 2023 for a brief citing six cases ChatGPT invented. Use it for structure and summarising your own material, never for sourcing.
Ask a model for a report and you get something that reads like a report immediately: headings, an executive summary, a measured tone, a conclusion. It is genuinely impressive and almost entirely beside the point, because fluency was never the hard part of a report. Being right is.
That gap is where AI report writing goes wrong, and it goes wrong in a specific, predictable way that is worth understanding before you file anything.
The citation problem
In June 2023 a federal court in New York sanctioned two lawyers $5,000 for a brief that cited six judicial decisions. None of the six existed. ChatGPT had produced them, and the attorney had used it as though it were a search engine.
The detail that matters most is what happened next. When the court could not find the cases and asked the lawyers to produce them, they did not withdraw the brief — they submitted excerpts of the opinions. Those excerpts were fabricated too. The model, asked to supply a thing that did not exist, cheerfully supplied it twice.
That is the failure mode in a sentence. A language model does not have a concept of "I could not find this". Asked for a source, it produces something source-shaped: a plausible author, a plausible year, a plausible title, formatted correctly. It is not lying, in the sense of knowing better. It is doing exactly what it does, which is produce likely-looking text.
Reports run on citations and figures. Which makes this the single thing to design your process around.
What AI report writing is genuinely good at
Everything that does not require knowing a fact.
Structure. Give it your topic and audience, and ask for a section outline. Arguing with an outline is faster than producing one.
Summarising material you supply. This is the strongest use by distance. Paste your own data, notes, transcripts or source documents, and ask for a summary, an executive summary, or the three findings a reader will care about. The facts come from you; the model only compresses. Nothing is invented because nothing needed to be retrieved.
The executive summary, last. Write the report, then ask for a 150-word summary of what you wrote. Models are reliably good at this and people are reliably bad at it, because by then you are too close to the material.
Tightening. Paste a paragraph, ask for it shorter. Mechanical, dull, and the output is easy to check because you already know what it should say.
Tone. Turning blunt internal notes into something appropriate for a board or a client, without losing the content.
Notice the pattern: in every one of these, the facts originate with you and the model rearranges them. That is the safe half of the job.
Numbers deserve their own warning
Models do arithmetic unreliably, and more importantly they do it invisibly. A wrong total in a paragraph of prose looks exactly like a right one — there is no error, no flag, nothing to notice. In a spreadsheet you might catch an odd-looking figure; in a sentence you will not.
Anything that will be read as a number — percentages, growth rates, totals, market sizes — should come from a source you can point at, and the arithmetic should be done somewhere that shows its working. If a model offers you a market-size figure, treat it as a hint about what to go and look up, never as the figure.
A process that actually holds up
- Gather sources first, yourself, before opening a chat window. This is the step that prevents every problem below it.
- Paste them in and ask for a structure, or for a summary of what you have gathered.
- Write the analysis yourself. This is your actual contribution and the reason the report has your name on it.
- Ask for an executive summary of your finished draft.
- Verify every citation and every number against the source, individually. Not a sample — every one. The fabrications look identical to the real ones, so spot-checking does not work.
Step one and step five are the whole discipline. If you supply the facts and check the facts, the model cannot introduce one.
Which tools
Long reports need context room — enough that the model can hold your source material and your draft at once. Claude is the usual pick for long documents for that reason, and its context window has been expanded specifically for long-document work. ChatGPT handles the same jobs and is better at ad-hoc data questions when you upload a file.
Notion AI is worth considering if the report lives alongside your notes anyway, since the summarising use works best when your material is already in the same place.
Our best AI writing tools list covers the wider category, with a free-only version, and the prompt engineering basics guide helps if your drafts keep coming back generic. For the related job of shaping a document about yourself rather than about data, see AI resume writing.
FAQ
Can AI write a full report for me?
It can produce something that reads like one, which is not the same thing. Any report that depends on real sources or real figures needs you to supply and verify both, because the model will invent them when it cannot find them — and the invented versions are formatted exactly like the real ones.
Does AI make up references?
Yes, routinely. A New York court sanctioned two lawyers $5,000 in June 2023 over a brief citing six non-existent cases produced by ChatGPT, and when asked to produce the cases they submitted fabricated excerpts. Check every citation against the actual source, individually.
What is the safest way to use AI for a report?
Supply the facts yourself. Paste your own sources, notes or data and ask for summarising, structuring and tightening. When the material originates with you, the model has nothing to invent, which removes the entire failure mode.
Can AI do the calculations in my report?
Treat anything numeric as unverified. Models get arithmetic wrong and the error is invisible in prose — a wrong total reads exactly like a right one. Calculate somewhere that shows its working and paste the result in.
Which AI is best for long reports?
Claude is generally preferred for long documents because of its context window, which lets it hold your source material and draft together. ChatGPT is stronger for questions about an uploaded data file. For most report work the difference matters less than whether you supplied the facts.
Related tools
Claude
Anthropic's AI assistant known for careful reasoning and long context.
ChatGPT
The most widely used AI chatbot for writing, coding, and research.
Notion AI
AI writing and Q&A assistant built into the Notion workspace.