Claude Chat proposed that the nineteenth-century scholar Rahmatullah Kairanawi had originated a particular argument. The argument is that the Greek paraklētos in John’s gospel was originally periklytos, which would make it a rough match for the name Ahmad. It sounded right. Correct period, correct figure, correct kind of claim. It shaped three rounds of research prompts before I confirmed it.
When I downloaded Kairanawi’s argument in the original Arabic and had my Claude OCR worker transcribe it, the argument had actually been proposed by the English orientalist George Sale in 1734, and Kairanawi had examined it and refused it.
Claude had proposed the attribution as a hypothesis, with a hedge. It wasn’t a fabricated citation. It was a plausible guess. Claude stopped flagging it as a guess in later turns, and it hardened into a premise because nobody tested it. What let it sit unchallenged is that Kairanawi’s refusal lives in the sixth book of a work whose standard English translation stops at the fifth. Nothing in English contradicts the claim, because nothing in English contains the passage.
Anyone who works with primary sources knows this shape of problem. A will summarised in a calendar entry rather than transcribed. A clinical finding that reaches you through a review article instead of the trial. The error is small, it sits in the part nobody translated, and everything downstream inherits it.
A wrong footnote costs you a correction. A wrong hypothesis costs you weeks.
Seven errors turned up during this project. None of them reached publication, because all seven got caught at the gate the process specifies for catching them.
The test case
I have been having some genuinely interesting conversations with a small group of very diverse college-age friends about the historical facts and other proofs that major religions use. So I decided to apply my current AI assisted research methodology to it.
Not a comparison of religions, and not a defence or attack on any of them. Just the same discipline I would apply to any other research question. A claim is made, check it against the evidence, and see whether it survives its own stated terms.
I am an independent researcher working across several domains. The framework is domain-agnostic, and it’s currently running in genealogy, healthcare and elsewhere.
What made this a useful test wasn’t that the subject was unfamiliar. It was three properties of the material. The topic is contested, with adversarial literatures on both sides, so secondary sources lean in predictable directions. The primary sources are in a script I cannot read. And it is a field where a fabricated citation looks exactly like a real one: plausible author, plausible title, plausible page range, plausible argument.
If the framework was going to fall over, it had good conditions to fall over in.
What transferred, and what didn’t
The governance transferred cleanly. Staged development, internal audit at release candidate, external verification on platforms that didn’t produce the work, primary-source checking, and honest disclosure of any gate left open.
The research instrument did not transfer at all. In genealogy I build transmission graphs and test directional sequence. Here I built a claim-testing rubric with rated domains: attestation, dating security, what premises a reader has to grant, how strong the best surviving counter-argument is. Completely different tools, same workflow.
That distinction is the useful part. The governance is domain-agnostic. The instruments it governs are domain-specific. Anyone hoping for a single method that answers questions in genealogy, healthcare and comparative religion is hoping for the wrong thing. What travels is the discipline around the method, not the method itself.
The staging comes from twenty-odd years of software delivery in a highly regulated industry, which I’ve written about in Governance by Design and am developing into a peer-reviewed article. Early drafts are development and unit test: internal iteration, adversarial review inside the same session, no external audit. External QA starts when the argument stops moving between drafts, which on this project was version six. Sending an unstable draft out for verification wastes audit cycles on text that will change structurally before anyone reads it again.
The errors, by type and by cost
Fabricated attribution. A claim attached to a real author and a real book, except the source doesn’t make it. This is the one unit testing cannot catch, because the reviewing model has no independent access to the artifact. It sits there among the correct citations looking exactly like them.
Fabricated bibliographic detail. Invented first names for two real authors, in a citation that was otherwise fine. Neither external platform caught it. A parallel session did. Slightly uncomfortable lesson: the QA gate isn’t one external pass, it’s however many independent readings you can afford.
Truncated quotation. An interpretive option presented as an author’s settled position. I’ll come back to this one, because it turned into something more interesting than a correction.
An unchecked hypothesis. The error from the opening, and the expensive one by a distance. Everybody verifies citations, because everybody knows citations can be wrong. A hypothesis that a platform floats mid-conversation just gets absorbed into the direction of the work, and a bad one takes everything downstream with it. Same gate, applied earlier.
On the tools, for the record: seven errors in total. Perplexity Pro Academic found four. SuperGrok found one, and it was methodological rather than factual, spotting that I had retrofitted a claim about predictive success onto research that had merely noticed a pattern after the fact. A parallel session found the invented first names. The Kairanawi hypothesis fell to all three searches converging on the same null. In the same round it caught the retrofit, SuperGrok’s most confident assertion was unsourced and wrong.
I’m reporting the asymmetry rather than tidying it up, because “the external audit caught the errors” is less useful than what actually happened. I suspect the split is structural rather than luck. Perplexity Pro Academic is search-first with scholarly connectors wired in, and every error it caught was a citation or source error. SuperGrok caught a reasoning error, which is what an adversarial engine is for. Neither found the other’s, and the one both missed took a third reader entirely.
The control that wasn’t in my framework
Cross-platform verification works beautifully on fabrication. A fabricated citation fails differently on a differently-trained model, so the second platform catches what the first invented.
It did nothing at all for summary-level drift, at least here, and this project made that painfully clear.
Perplexity Pro Academic and SuperGrok both answered questions about a primary text by reading summaries of it, while the text itself sat one click away. One couldn’t extract a PDF that opened cleanly for another tool. A search tool told me a document appeared in only two places online while I was looking at a third instance sitting in my own files.
What worked, over and over, was supplying the document. Not a report of the source. The source. The moment I put page images of the Arabic critical edition in front of a model, a question that had eaten three rounds of searching resolved in one exchange.
So the framework formalized an intuitive rule. Any load-bearing claim about what a text says requires someone to open the text. Cross-platform verification is necessary, and it has to validate against the original document wherever possible.
How hard that is varies enormously. Colonial probate records are straightforward. Esoteric nineteenth-century Arabic is not, and neither is Latin in secretary hand or a Norwegian township register. The answer isn’t a specialist on call for every language you meet. It’s the same chain you’d run anywhere else: OCR with confidence scoring, independent transcriptions compared against each other, published translations used as a check, and the original text quoted in full so a reader can verify it themselves. That’s more auditable than a private consultation, because it leaves a trail anyone can follow.
Two findings, not corrections
The standing complaint about AI-assisted research is that it just rearranges what already exists. Two results from this project push back on that, and both came from chasing errors instead of quietly patching them.
The truncation that wasn’t the AI’s fault. The system quoted the medieval commentator Fakhr al-Din al-Razi asserting that certain biblical narratives matched the Qur’an exactly, without any difference at all. Verification against the Arabic showed the clause was the second of four interpretive options al-Razi lists. Completely standard exegetical practice. Not a settled claim at all.
But the AI hadn’t invented the truncation. That is how the passage circulates in the polemical literature, and it had been picked up faithfully. Worth noting that this is the second time a draft of mine arguing for citation discipline has turned out to contain the exact fault it was arguing against. I am starting to think the failure mode clusters around whatever you happen to be writing about.
Then, while checking something unrelated, the same structure showed up running the other way. Raymond Brown, the Catholic biblical scholar, opens an appendix by noting that various scholars have doubted a particular identification, and then spends nine pages testing that view and rejecting it. The version circulating in Muslim apologetics stops at the doubt. It deletes his next sentence, the one announcing that he is about to test the claim. Some versions add a sentence he never wrote, promising that proof is coming.
Two cases, opposite confessional directions, one technique. Take an expert from the other tradition, strip off the frame that makes the statement provisional, present the result as testimony against his own side.
Two cases are not a pattern, and I’m not going to pretend otherwise. What they are is a hypothesis worth testing against a corpus, and it exists only because somebody followed an error back to its source instead of fixing it and moving on.
The passage nobody had read. Verifying that opening hypothesis meant getting page images from the Arabic critical edition. What they showed was better than the correction. Kairanawi had received the argument from a named missionary tract, printed in Calcutta in a specified year, published by the very people the argument was going to be used against. He credited it as plausible. Then he set it aside as not good enough for his purposes and built his case on different ground entirely.
That discussion sits in the sixth book of his work. The standard English translation contains the first five. Four independent checks, three searches across six languages plus a recent full-length study of the book by a specialist, found nobody engaging that passage. One reference volume remains unchecked behind a paywall, which is the caveat that belongs with the finding rather than after it.
I’m not claiming a major discovery here. I’m claiming a verification chain can produce findings and not just catch mistakes, which is a different thing from rearrangement.
What “it worked” actually means
Seven errors. All caught before publication, at the gate the process specifies for catching them.
The value was never error prevention. No matter how good your prompt, it doesn’t stop a system from generating a plausible fabrication, which is the argument I made at more length in Four Platforms, One Standard. Prompting is the first discipline, not the last defence. The value is that errors surface while they are still cheap. In a draft, not in print. In a footnote, not in a conclusion resting on one.
One gate is still open. That paywalled reference volume is the last place a contrary finding could be hiding, and it’s disclosed in the article rather than quietly closed, which is the other half of the discipline.
The article is unpublished anyway, since I have two remaining religions to finish plus a comparative summary. But the seven errors are already dealt with, and that is what a working method looks like. Not one that prevents mistakes. One that makes them cheap.
Richard E. Rudd is an independent researcher working across multiple fields. He spent more than twenty years as an IT Product Owner and Program Manager at large multinational financial services and insurance firms before applying those governance disciplines to AI-assisted research. This post describes one project run under that framework; the article it discusses is in preparation.