Blog

  • What My Research Framework Caught: Seven Errors, None Published

    Claude Chat proposed that the nineteenth-century scholar Rahmatullah Kairanawi had originated a particular argument. The argument is that the Greek paraklētos in John’s gospel was originally periklytos, which would make it a rough match for the name Ahmad. It sounded right. Correct period, correct figure, correct kind of claim. It shaped three rounds of research prompts before I confirmed it.

    When I downloaded Kairanawi’s argument in the original Arabic and had my Claude OCR worker transcribe it, the argument had actually been proposed by the English orientalist George Sale in 1734, and Kairanawi had examined it and refused it.

    Claude had proposed the attribution as a hypothesis, with a hedge. It wasn’t a fabricated citation. It was a plausible guess. Claude stopped flagging it as a guess in later turns, and it hardened into a premise because nobody tested it. What let it sit unchallenged is that Kairanawi’s refusal lives in the sixth book of a work whose standard English translation stops at the fifth. Nothing in English contradicts the claim, because nothing in English contains the passage.

    Anyone who works with primary sources knows this shape of problem. A will summarised in a calendar entry rather than transcribed. A clinical finding that reaches you through a review article instead of the trial. The error is small, it sits in the part nobody translated, and everything downstream inherits it.

    A wrong footnote costs you a correction. A wrong hypothesis costs you weeks.

    Seven errors turned up during this project. None of them reached publication, because all seven got caught at the gate the process specifies for catching them.

    The test case

    I have been having some genuinely interesting conversations with a small group of very diverse college-age friends about the historical facts and other proofs that major religions use. So I decided to apply my current AI assisted research methodology to it.

    Not a comparison of religions, and not a defence or attack on any of them. Just the same discipline I would apply to any other research question. A claim is made, check it against the evidence, and see whether it survives its own stated terms.

    I am an independent researcher working across several domains. The framework is domain-agnostic, and it’s currently running in genealogy, healthcare and elsewhere.

    What made this a useful test wasn’t that the subject was unfamiliar. It was three properties of the material. The topic is contested, with adversarial literatures on both sides, so secondary sources lean in predictable directions. The primary sources are in a script I cannot read. And it is a field where a fabricated citation looks exactly like a real one: plausible author, plausible title, plausible page range, plausible argument.

    If the framework was going to fall over, it had good conditions to fall over in.

    What transferred, and what didn’t

    The governance transferred cleanly. Staged development, internal audit at release candidate, external verification on platforms that didn’t produce the work, primary-source checking, and honest disclosure of any gate left open.

    The research instrument did not transfer at all. In genealogy I build transmission graphs and test directional sequence. Here I built a claim-testing rubric with rated domains: attestation, dating security, what premises a reader has to grant, how strong the best surviving counter-argument is. Completely different tools, same workflow.

    That distinction is the useful part. The governance is domain-agnostic. The instruments it governs are domain-specific. Anyone hoping for a single method that answers questions in genealogy, healthcare and comparative religion is hoping for the wrong thing. What travels is the discipline around the method, not the method itself.

    The staging comes from twenty-odd years of software delivery in a highly regulated industry, which I’ve written about in Governance by Design and am developing into a peer-reviewed article. Early drafts are development and unit test: internal iteration, adversarial review inside the same session, no external audit. External QA starts when the argument stops moving between drafts, which on this project was version six. Sending an unstable draft out for verification wastes audit cycles on text that will change structurally before anyone reads it again.

    The errors, by type and by cost

    Fabricated attribution. A claim attached to a real author and a real book, except the source doesn’t make it. This is the one unit testing cannot catch, because the reviewing model has no independent access to the artifact. It sits there among the correct citations looking exactly like them.

    Fabricated bibliographic detail. Invented first names for two real authors, in a citation that was otherwise fine. Neither external platform caught it. A parallel session did. Slightly uncomfortable lesson: the QA gate isn’t one external pass, it’s however many independent readings you can afford.

    Truncated quotation. An interpretive option presented as an author’s settled position. I’ll come back to this one, because it turned into something more interesting than a correction.

    An unchecked hypothesis. The error from the opening, and the expensive one by a distance. Everybody verifies citations, because everybody knows citations can be wrong. A hypothesis that a platform floats mid-conversation just gets absorbed into the direction of the work, and a bad one takes everything downstream with it. Same gate, applied earlier.

    On the tools, for the record: seven errors in total. Perplexity Pro Academic found four. SuperGrok found one, and it was methodological rather than factual, spotting that I had retrofitted a claim about predictive success onto research that had merely noticed a pattern after the fact. A parallel session found the invented first names. The Kairanawi hypothesis fell to all three searches converging on the same null. In the same round it caught the retrofit, SuperGrok’s most confident assertion was unsourced and wrong.

    I’m reporting the asymmetry rather than tidying it up, because “the external audit caught the errors” is less useful than what actually happened. I suspect the split is structural rather than luck. Perplexity Pro Academic is search-first with scholarly connectors wired in, and every error it caught was a citation or source error. SuperGrok caught a reasoning error, which is what an adversarial engine is for. Neither found the other’s, and the one both missed took a third reader entirely.

    The control that wasn’t in my framework

    Cross-platform verification works beautifully on fabrication. A fabricated citation fails differently on a differently-trained model, so the second platform catches what the first invented.

    It did nothing at all for summary-level drift, at least here, and this project made that painfully clear.

    Perplexity Pro Academic and SuperGrok both answered questions about a primary text by reading summaries of it, while the text itself sat one click away. One couldn’t extract a PDF that opened cleanly for another tool. A search tool told me a document appeared in only two places online while I was looking at a third instance sitting in my own files.

    What worked, over and over, was supplying the document. Not a report of the source. The source. The moment I put page images of the Arabic critical edition in front of a model, a question that had eaten three rounds of searching resolved in one exchange.

    So the framework formalized an intuitive rule. Any load-bearing claim about what a text says requires someone to open the text. Cross-platform verification is necessary, and it has to validate against the original document wherever possible.

    How hard that is varies enormously. Colonial probate records are straightforward. Esoteric nineteenth-century Arabic is not, and neither is Latin in secretary hand or a Norwegian township register. The answer isn’t a specialist on call for every language you meet. It’s the same chain you’d run anywhere else: OCR with confidence scoring, independent transcriptions compared against each other, published translations used as a check, and the original text quoted in full so a reader can verify it themselves. That’s more auditable than a private consultation, because it leaves a trail anyone can follow.

    Two findings, not corrections

    The standing complaint about AI-assisted research is that it just rearranges what already exists. Two results from this project push back on that, and both came from chasing errors instead of quietly patching them.

    The truncation that wasn’t the AI’s fault. The system quoted the medieval commentator Fakhr al-Din al-Razi asserting that certain biblical narratives matched the Qur’an exactly, without any difference at all. Verification against the Arabic showed the clause was the second of four interpretive options al-Razi lists. Completely standard exegetical practice. Not a settled claim at all.

    But the AI hadn’t invented the truncation. That is how the passage circulates in the polemical literature, and it had been picked up faithfully. Worth noting that this is the second time a draft of mine arguing for citation discipline has turned out to contain the exact fault it was arguing against. I am starting to think the failure mode clusters around whatever you happen to be writing about.

    Then, while checking something unrelated, the same structure showed up running the other way. Raymond Brown, the Catholic biblical scholar, opens an appendix by noting that various scholars have doubted a particular identification, and then spends nine pages testing that view and rejecting it. The version circulating in Muslim apologetics stops at the doubt. It deletes his next sentence, the one announcing that he is about to test the claim. Some versions add a sentence he never wrote, promising that proof is coming.

    Two cases, opposite confessional directions, one technique. Take an expert from the other tradition, strip off the frame that makes the statement provisional, present the result as testimony against his own side.

    Two cases are not a pattern, and I’m not going to pretend otherwise. What they are is a hypothesis worth testing against a corpus, and it exists only because somebody followed an error back to its source instead of fixing it and moving on.

    The passage nobody had read. Verifying that opening hypothesis meant getting page images from the Arabic critical edition. What they showed was better than the correction. Kairanawi had received the argument from a named missionary tract, printed in Calcutta in a specified year, published by the very people the argument was going to be used against. He credited it as plausible. Then he set it aside as not good enough for his purposes and built his case on different ground entirely.

    That discussion sits in the sixth book of his work. The standard English translation contains the first five. Four independent checks, three searches across six languages plus a recent full-length study of the book by a specialist, found nobody engaging that passage. One reference volume remains unchecked behind a paywall, which is the caveat that belongs with the finding rather than after it.

    I’m not claiming a major discovery here. I’m claiming a verification chain can produce findings and not just catch mistakes, which is a different thing from rearrangement.

    What “it worked” actually means

    Seven errors. All caught before publication, at the gate the process specifies for catching them.

    The value was never error prevention. No matter how good your prompt, it doesn’t stop a system from generating a plausible fabrication, which is the argument I made at more length in Four Platforms, One Standard. Prompting is the first discipline, not the last defence. The value is that errors surface while they are still cheap. In a draft, not in print. In a footnote, not in a conclusion resting on one.

    One gate is still open. That paywalled reference volume is the last place a contrary finding could be hiding, and it’s disclosed in the article rather than quietly closed, which is the other half of the discipline.

    The article is unpublished anyway, since I have two remaining religions to finish plus a comparative summary. But the seven errors are already dealt with, and that is what a working method looks like. Not one that prevents mistakes. One that makes them cheap.

    Richard E. Rudd is an independent researcher working across multiple fields. He spent more than twenty years as an IT Product Owner and Program Manager at large multinational financial services and insurance firms before applying those governance disciplines to AI-assisted research. This post describes one project run under that framework; the article it discusses is in preparation.

  • Governance by Design: What Twenty Years in Regulated IT Taught Me About AI-Assisted Research

    Why should anyone care which AI tools I used, as long as the work meets the standard?

    I think the honest answer is: they shouldn’t. Nobody asks an auditor what calculator they used. Nobody asks a developer which IDE they coded in. The deliverable either satisfies the standard or it doesn’t. The tools are implementation details, disclosed for transparency — not because they require justification.

    But that answer only holds if something guarantees the standard was met. In regulated software development, that something has a name. It’s called governance, and it is the most underrated idea in the current conversation about AI and research.

    The Question Nobody Is Asking

    The AI-and-research conversation is currently stuck on two questions. Which tool is best? And is it ethical to use one at all?

    Both skip past the question that actually determines whether the work is trustworthy: what structure ensures the output meets the standard?

    That is a governance question. And governance is not an unsolved problem. It was solved decades ago, in industries where getting it wrong is expensive.

    I spent more than twenty years as an IT Product Owner and Program Manager at large multinational financial services and insurance firms, building business-critical systems under regulatory scrutiny. In that world, you cannot ship code because the developer says it works. The developer’s confidence is not evidence. What makes a deliverable trustworthy is never the skill of any individual — it is the structure built around them.

    Separation of duties. Independent QA. Code review by someone who did not write the code. Audit trails. Change control. Defect logging. Testing rigor scaled to risk.

    None of that assumes the developer is careless. It assumes the developer is human — subject to blind spots, overconfidence, and investment in their own work.

    AI has precisely the same failure modes, in sharper form. It is fluent, confident, eager to please, and — this is the part that should worry us — its errors are indistinguishable in appearance from its successes. A fabricated citation is formatted exactly like a real one.

    Which means the governance pattern transfers. Not perfectly — and the imperfections matter, so I’ll be specific about them. But the core structure transfers more directly than most analogies in this conversation, and it transfers because the underlying problem is the same: how do you trust output from a fluent, confident, self-interested producer?

    The Load-Bearing Principle

    Of everything I carried across from regulated development, one rule does most of the work:

    The platform that produced a finding cannot be the sole platform that verifies it.

    This is separation of duties. In financial auditing, the team that prepares the statements cannot be the team that audits them — not because they are dishonest, but because they are invested. In software, the developer who wrote the code does not sign off as the sole reviewer of their own work.

    Applied to AI-assisted research, the rule becomes: any platform can do QA. No platform does QA on its own work.

    That single constraint dissolves most of what people are actually afraid of when they worry about AI in research. Sycophancy is structurally defeated, because the auditing platform has no stake in the analyzing platform’s conclusion. Hallucination is substantially caught, because the same fabrication rarely surfaces on two independently trained models.[1] Confirmation bias is blunted, because the auditor receives the research question and the sources — not the conclusion it is meant to confirm.

    The rule is structural, not preferential. It does not depend on any model being good.

    Where the transfer is not clean

    I want to be precise about the disanalogies, because a careful reader will find them and I would rather name them first.

    In regulated software, the independent reviewer is a credentialed professional, legally accountable, operating under compliance regimes with real enforcement. A second AI platform has none of that. It is a different commercial product, not an independent professional bound by duty. So the transfer is structural, not institutional — and the difference is worth stating plainly.

    What survives the difference is the mechanism. A human auditor’s credential guarantees competence. An AI auditor’s independence guarantees something narrower but still valuable: a different error surface. The second platform was trained on different data with a different architecture, so it fails in different places. That is not the full institutional protection of regulated auditing. It is a real, structural reduction in correlated error — and for research quality control, that is most of what you need. The rest, as I’ll argue at the end, stays human.

    One further limit belongs in the same honest accounting: today’s frontier models share substantial training-corpus overlap, so the diversity of error surfaces is real but not total. Where a mistake is baked into the common substrate of the web these models learned from, two platforms can still agree and still be wrong. Cross-platform verification reduces correlated error; it does not eliminate it. That residual is one more reason the final judgment stays with the researcher, not the machines.

    The Field Is Converging on This — From Several Directions

    The interesting thing about multi-model verification is that it is being independently rediscovered right now, from several directions at once.

    Andrej Karpathy built an “LLM Council” that routes a question to several frontier models, has them review one another’s answers anonymously, and appoints a chairman model to synthesize (he called it a weekend project, but the architecture makes a serious point).[2] The platform vendors have begun to build the pattern in — some agentic modes now spawn parallel sub-agents that check a result before returning it, though most current implementations are orchestration for speed rather than structured cross-verification.[3] And Steve Little, the National Genealogical Society’s AI Program Director, has published a clean demonstration of it: put two different reasoning engines at the same workbench, give them the same evidence and the same method, and have them read independently before either sees the other’s answer. When they disagree, the disagreement is the signal — a question surfaced that one confident answer would have hidden.[4]

    Three different arrivals at the same underlying instinct: one model is not enough, and the way to trust the output is to make independent readers check each other. That convergence is a good sign the instinct is sound.

    It also lets me draw a distinction the convergence tends to blur — one that matters for how you actually build the thing.

    Start with the multi-agent modes. A council of agents inside a single model is still one training corpus, one set of corporate response policies, one architecture. It catches inconsistency — places where the model contradicts itself — but it cannot catch a blind spot the whole model shares. If the training data carries a systematic error, every agent inherits it, and unanimous agreement among them means only that they are all wrong together.[5] Cross-agent verification within one model reduces variance. Cross-platform verification across independently trained models is what actually catches shared blind spots. Karpathy’s own design implies this — he built a council of genuinely different models, not one model talking to itself.

    Now the sharper design question, the one Steve Little raises directly. If you do use different platforms, should they hold fixed roles? His worry is vendor mythology — the comfortable story that this model is the careful analyst and that one is the creative writer, as though these were personalities rather than products. His remedy is role reversal: today’s extractor becomes tomorrow’s skeptic, so no model hardens into an unquestionable authority. As he puts it, “a model personality is merely a story we may be tempted to tell.”[6]

    He is right, and the mechanism is a good one. Role reversal does something a fixed audit function doesn’t: it keeps the researcher fluent across every model and prevents any one of them from becoming a black box you stop checking. For the task his article addresses — two engines reading the same evidence to answer the same question — it’s the right call. But it isn’t the whole picture.

    Where I’d extend the picture is one layer up. The danger in vendor mythology is not specialization itself. It is deference — trusting the designated expert without checking the work. Those two things are separable, and regulated software separates them every day. You do not cure developer overconfidence by rotating developers into QA on a schedule; you cure it with a QA function that reviews everything, permanently. Fixed roles, zero self-review.

    That distinction matters because the differences between platforms are real — and some of them are not stories at all.

    Some differences are behavioural: a model’s characteristic verbosity, its hedging, its appetite for speculation. These emerge from training data and corporate response policy, which is why the quirks are consistent rather than random. They are also largely mitigable through prompting — a well-built prompt flattens them, and Steve’s method depends on exactly that flattening. On this point the vendor-mythology worry is fair.

    But other differences are structural. A search-first architecture with mandatory citations wired into a scholarly index will surface and verify literature that a general-purpose model simply cannot reach. Access to data corpora a competitor cannot license at any price. Agentic tooling that operates directly on files and archives. No prompt closes those gaps, because they are not behavioural — they are plumbing. Assigning a scholarly-literature task to the platform actually engineered for scholarly literature is not mythology. It is matching the work to the tool.

    So the full picture has two layers. When platforms read the same evidence to answer the same question — adjudication — roles should be symmetric and reversible, precisely as Little argues. When platforms occupy different stages of a pipeline — discovery, retrieval, deep analysis, literature verification, drafting — differentiated assignment is justified by real structural advantages. Both layers live under one rule: no platform grades its own homework.

    I expect the structural gaps to narrow as the platforms advance, and when they do, the pipeline layer will look more like the adjudication layer — more interchangeable, more reversible. We are not there yet. And a framework organized by task rather than by model absorbs that convergence without breaking: as the differences shrink, you simply reverse and rotate more freely, under the same governance. The one refinement I’d add even now is that structural advantages shift over twelve to eighteen months, so even fixed assignments deserve periodic re-benchmarking. Fixed where the structural edge is real and current; reversible where it isn’t; never self-reviewing anywhere.

    What Else Transfers

    Errors are the audit trail, not an embarrassment. In regulated development, defects are logged, not buried. Every error caught gets documented: what was claimed, by which platform, what flagged it, what the primary source actually showed. A methodology that reports no errors is not rigorous. It is untested. My own errors — including one in a draft arguing for citation verification, which itself contained a misattributed citation, caught by the audit step — are the strongest evidence I have that the process works.

    Locked decisions. Once a conclusion clears full verification, it is locked. A later session cannot quietly revert it. That is change control. You do not roll back a production release because someone had second thoughts on a Tuesday.

    Governance scales to risk. Nobody applies the same rigour to a two-week bug fix and a two-year platform migration. Verifying a single reported figure needs a quick cross-check between two models. A contested conclusion bound for peer review needs the full apparatus.

    The human role is non-delegable. In SDLC terms, the researcher is the Product Owner. They define the question, accept or reject the deliverable, and remain accountable for whether it meets the standard. The AI platforms are specialized developers. They do not write their own requirements, and they do not sign their own acceptance. AI enables, the methodology governs, the researcher decides — and the last of those three can never be delegated.

    Why This Isn’t About Any One Field

    I am an independent researcher, and I work across several fields. The framework I am describing did not come out of any of them. It came out of regulated IT, and I carried it in.

    I have developed and tested it most fully in genealogy, for a specific reason: genealogy has an unusually clean published evidentiary standard. The Genealogical Proof Standard defines what a trustworthy conclusion looks like independently of who reached it or what tools they used. That makes it an ideal proving ground — there is an external yardstick, and the yardstick does not care about my methods. It is one of my current research areas, not my field, and that distinction is the whole point: nothing in the governance framework is genealogical.

    It is:

    •          Tool-agnostic in principle, specialized in practice. The framework does not require any particular platforms — only that you use at least two, and that none of them grades its own homework. But the best way to run it today exploits the specific structural advantages different platforms actually have. The framework will outlive the current lineup; the current lineup is still the sharpest way to implement it.

    •          Topic-agnostic. I am currently running it against medical research as well, where the producers are just as fluent and the failure modes just as costly.

    •          Standards-agnostic. Any field with a defined evidentiary standard can be plugged in: legal, clinical, archival, journalistic.

    •          Scalable. Two free-tier models for a simple question. The full apparatus for publication-track scholarship.

    The claim I am making is not “here is my workflow.” Workflows are personal, and they are obsolete within the year. The claim is that regulated industries already solved the problem of trusting output from fallible, confident, self-interested producers — and that the solution transfers to AI-assisted research more or less intact. Notably, the major AI-governance frameworks don’t yet make this transfer: they govern how organizations build and deploy AI systems, not how an individual researcher should structure AI-assisted inquiry against an evidentiary standard.[7]

    What Comes Next

    Two related papers are in preparation. The first develops the governance paradigm and its transfer into the digital humanities. The second applies it to a specific evidentiary standard, with worked case studies — including the failures, which are the interesting part. Preprint articles are expected to follow shortly.

    The tools will keep changing. The governance does not have to.

    Notes

    Richard E. Rudd is an independent researcher working across multiple fields. He spent more than twenty years as an IT Product Owner and Program Manager at large multinational financial services and insurance firms before applying those governance disciplines to AI-assisted research.

    [1] The empirical basis for this is developing quickly. See P. Verga et al., “Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models” (arXiv:2404.18796, 2024), which finds a panel of diverse models reduces intra-model bias relative to any single-model judge. Enterprise-scale studies in regulated sectors report substantial hallucination reduction from cross-platform verification.

    [2] Andrej Karpathy, “LLM Council,” github.com/karpathy/llm-council (2025). Three-stage architecture: parallel individual responses, anonymized peer review and ranking, and a chairman model that synthesizes. Karpathy described it as a weekend project; the design nonetheless illustrates the cross-model principle precisely.

    [3] As of mid-2026, OpenAI, Anthropic, and Google have shipped multi-agent SDKs and orchestration frameworks, and some “deep research” products deploy parallel sub-agents. Most current implementations orchestrate for parallelism rather than structured cross-verification; the verification pattern is more fully realized in open-source projects and dedicated verification layers.

    [4] Steve Little, “Trifecta: One Genealogist, Two AI Assistants, One Folder,” Vibe Genealogy (vibegenealogy.ai), 10 July 2026. The “vendor mythology” and role-reversal argument, and the workbench method of independent reads adjudicated against the record, are set out there in full. Little’s role as the National Genealogical Society’s AI Program Director is per the NGS announcement of the position (ngsgenealogy.org).

    [5] On the self-review limit, see the “self-correction blind spot” literature (e.g., “Self-Correction Bench,” arXiv:2507.02778, 2025), which finds that models fail to catch a majority of their own internal errors while reliably catching the identical errors when presented as external input. This is the empirical core of the case for cross-platform rather than same-model checking.

    [6] Steve Little, “Trifecta: One Genealogist, Two AI Assistants, One Folder,” Vibe Genealogy (vibegenealogy.ai), 10 July 2026. The “vendor mythology” and role-reversal argument, and the workbench method of independent reads adjudicated against the record, are set out there in full. Little’s role as the National Genealogical Society’s AI Program Director is per the NGS announcement of the position (ngsgenealogy.org).

    [7] The NIST AI Risk Management Framework and ISO/IEC 42001 both address separation of duties and independent oversight — but at the level of organizational governance of AI systems, not as a methodology for structuring an individual researcher’s AI-assisted work against an evidentiary standard. That specific application is the gap this work addresses.

  • Four Platforms, One Standard: Why Serious AI-Assisted Research Needs More Than a Better Prompt

    Every AI model gets things wrong. That’s not controversial — it’s well-documented. Chat GPT hallucinates confident citations to papers that don’t exist.[1] Claude can construct logically elegant arguments from flawed premises. Gemini sometimes conflates similar but distinct concepts. Perplexity occasionally returns outdated information as current fact. And the problem goes deeper than factual errors: researchers have documented that even semantically similar prompts can produce drastically different outputs from the same model, a fragility known as “prompt brittleness.”[2]

    In my own research — tracing nine armigerous ancestral lines through the clothier and gentry families of Tudor Kent — I’ve watched three different AI platforms confidently produce three different, mutually exclusive interpretations of the same sixteenth-century will. Each interpretation was internally consistent, well-reasoned, and wrong.

    The standard response to these problems is to write better prompts. Be more specific. Add constraints. Use chain-of-thought reasoning. Provide few-shot examples. And all of that genuinely helps — these are real advances, and researchers teaching prompt engineering at conferences and in workshops are doing valuable work. A well-structured prompt with clear context, role assignment, and explicit formatting constraints will produce measurably better output than a vague one.[3]

    But you’re still asking a single model to check its own work. You’ve improved the employee, but you haven’t addressed the fundamental problem: one perspective, one training dataset, one set of blind spots.

    The problem is worse than most users realize. Researchers at MIT CSAIL recently formalized what they call “delusional spiraling” — situations where extended conversations with a single AI chatbot lead users to high confidence in demonstrably false beliefs.[4] The Human Line Project has documented nearly 300 such cases, including an accountant who, after weeks of single-chatbot conversation, came to believe he was trapped in a false universe.[5] The MIT team’s formal model demonstrates something unsettling: even an idealized, perfectly rational user is vulnerable to this effect when interacting with a sycophantic model. And two commonly proposed mitigations — preventing the chatbot from hallucinating, and informing users about model sycophancy — do not eliminate it.[6]

    This is the single-model problem stated in its most extreme form. But milder versions of the same dynamic affect every serious research project conducted with a single AI platform: the model’s agreeable, confident, internally consistent output gradually shapes the researcher’s thinking, and there’s no independent check in the system.

    Beyond Single Models

    Andrej Karpathy proposed an elegant next step — the LLM council, where you query multiple models and synthesize their responses.[7] It’s a real improvement over single-model work, and for many use cases it’s sufficient. But there’s a limitation that matters for research: large language models are trained on overlapping internet-scale datasets. Because they share training data, they are likely to share blind spots — a form of correlated error that an LLM council may not catch, even when the individual models would each flag a different mistake in isolation.

    This limitation is real but relative, not absolute. Current frontier models do use different architectures, different fine-tuning approaches, and different post-training processes, which produces genuinely different error profiles on many tasks. The methodology treats multi-platform verification as an additional layer of defense, not a final one — just as the Genealogical Proof Standard never permits reliance on a single source, this framework never treats cross-platform agreement as a substitute for the primary archival record.

    None of this means that better prompting or LLM councils are wasted effort. Prompt engineering is the first discipline of serious AI-assisted research, and this methodology assumes that competence. Chain-of-thought reasoning, few-shot examples, role-assignment techniques — these are meaningful advances that make AI more useful for everyone. An LLM council that queries multiple models is better than querying one. The methodology I’m describing doesn’t replace these approaches. It builds on them. It starts from the premise that you’ve already written a good prompt for a capable model, and asks: what structural verification do you still need?

    The chess world offers a useful parallel for thinking about this progression. As Dr. Dominic Ng recently observed, chess is “30 years ahead of every other profession in dealing with AI.”[8] In 2005, two amateurs with laptops famously beat both grandmasters and supercomputers by combining human and machine input more effectively than either could operate alone.[9] It was the triumph of “human in the loop.” But by 2026, the dynamic has reversed: adding a human to a chess engine actually makes it play worse.[10] “Human in the loop” wasn’t a permanent solution — it was a transitional phase.

    But genealogical research is not chess. In chess, the outcome is computationally verifiable — there is a correct move, and a sufficiently powerful engine will find it. Genealogical conclusions require judgment about ambiguous, incomplete, and contradictory evidence that no AI can resolve for itself. The Genealogical Proof Standard codifies that requirement: its five elements demand not just evidence but reasoned human evaluation of evidence, including the resolution of conflicts that may have no algorithmic solution.[11] The lesson from chess isn’t that human judgment is obsolete. It’s that how humans and AI interact matters more than whether they interact. The value isn’t in being “in the loop” as a vague reassurance — it’s in designing the loop itself.

    When Standards Are at Stake

    In fields where research must meet a defined evidentiary standard — where conclusions aren’t just opinions but claims that can be independently verified — single-model AI assistance creates a specific and dangerous problem: it can produce work that looks rigorous while containing errors that the model itself cannot detect.

    The separation-of-duties principle in regulated industries offers a conceptual parallel. In financial auditing, the people who prepare the accounts are never the same people who audit them — a principle codified in standards like ISA 610 and SOX Section 404.[12] Pharmaceutical companies apply similar logic to drug discovery validation. The architectural principle is the same: no single system, human or artificial, should be trusted to validate its own output. Some regulated industries are actively developing independent AI validation workflows; the principle is well established even where the specific multi-platform practice is still emerging.

    Genealogy has its own published evidentiary standard — the Genealogical Proof Standard (GPS), a five-element framework that governs how conclusions are reached and documented in peer-reviewed genealogical scholarship.[13] The GPS requires reasonably exhaustive searches, complete and accurate source citations, skilled analysis of evidence, resolution of conflicting evidence, and soundly reasoned written conclusions. It doesn’t care whether your tools are digital or analog. It cares whether your process is defensible.

    The same structural question faces historians working with digitized archives, anthropologists analyzing field data, and social scientists conducting qualitative research: how do you integrate AI into a research process that must meet a defined evidentiary or methodological standard?

    A Field Ready for Governance

    The research community is already moving toward an answer. In genealogy, the Coalition for Responsible AI in Genealogy (CRAIGEN) has published guiding principles for ethical AI use, emphasizing accuracy, disclosure, and compliance with existing standards.[14] Steve Little, the National Genealogical Society’s AI Program Director, has led workshops and webinars on responsible AI adoption, including GPS-compliant narrative writing with AI assistance — work that has helped establish the vocabulary the community uses to discuss AI and evidentiary rigor. In a recent cross-platform test, Little fed the same 1909 newspaper article to both Claude and ChatGPT and got meaningfully different results: one platform extracted 34 individuals connected by family relationships, the other extracted 55 names including unlinked mentions — same source, same prompt, different architectural decisions about what constitutes genealogically relevant data.[15] Brian’s Ancestors and Algorithms podcast has recently explored GPS-mapped multi-platform workflows, assigning specific AI tools to specific GPS elements based on functional strengths — a practical demonstration that researchers are independently arriving at multi-platform architectures.[16]

    In adjacent fields, multi-agent AI pipelines for fact verification and bias reduction are emerging in clinical research and journalism, and university programs are beginning to teach cross-platform AI output verification as a core research skill.[17] The tools are available. The community recognizes the need for governance. What’s missing is a formalized framework — one that goes beyond best practices and workflow tips to define roles, verification protocols, and documentation requirements that make AI-assisted research as defensible as traditional scholarship.

    A Governance Methodology, Not a Tool Review

    Over the past year, I’ve been developing and field-testing exactly that: a structured multi-platform AI verification methodology for standards-governed research, demonstrated first in genealogy because that’s my current research domain, but designed to be portable to any field with a published evidentiary standard. The framework assigns different AI platforms to specialized roles — deep analysis, project orchestration, adversarial auditing, independent cross-verification — and governs the entire process through the five elements of the GPS.

    This is not “ask another chatbot.” It is a role-based verification system with locked decisions, confidence labels, and an audit trail. The methodology defines specific platform roles, establishes locked analytical decisions that prevent later sessions from reverting confirmed conclusions, enforces temporal rules about what kinds of evidence can upgrade what kinds of claims, and documents errors and corrections as rigorously as successes. It’s the difference between “I got a second opinion” and a structured audit process with defined responsibilities and a paper trail.

    The methodology rests on two pillars.

    Platform Specialization. Each AI platform has measurable strengths. Independent benchmarks — from Stanford’s HELM evaluation to LMSYS’s Chatbot Arena — confirm that no single model leads across all task categories.[18] Rather than using one platform for everything, the methodology assigns roles based on demonstrated capability — the same way a research team assigns tasks based on expertise. If you were hiring researchers, you wouldn’t hire five people with the same degree from the same university. You’d want a data analyst, a subject-matter expert, a skeptic whose job is to find holes, and a project manager. The methodology applies the same logic to AI platforms.

    Multi-platform architecture also unlocks capabilities that single-platform work cannot access. Research on automated prompt optimization — including Google DeepMind’s OPRO framework and Zhou et al.’s work demonstrating that LLMs are “human-level prompt engineers” — has shown that AI models can generate prompts that match or exceed human-crafted ones.[19] In a multi-platform methodology, this means one platform can generate queries optimized for another platform’s specific strengths — a form of inter-platform collaboration that builds on human prompt engineering skill rather than replacing it.[20]

    Cross-Platform Verification. No conclusion stands unless it has been independently examined by at least two platforms with different training data and different architectures. This is the digital equivalent of what the GPS already requires: no responsible genealogist trusts a single source for a critical conclusion.[21] The methodology extends that principle to the AI tools themselves. Crucially, this is the structural defense against the “delusional spiraling” that the MIT team documented — by breaking the single-model feedback loop, cross-platform verification introduces the independent check that single-model interaction inherently lacks.

    The framework is meant to be adopted in stages: start with two platforms and a no-single-source rule, then add roles and controls as the project grows. Many AI platforms offer free tiers; a practical starting configuration — one analytical hub and one adversarial research engine, both at paid tiers runs about $40 per month and delivers a meaningful improvement in both capability and verification rigor over any single-platform workflow. It maps to all five GPS elements, documenting not just what the AI contributed but what it got wrong and how the errors were caught.

    The key distinction is important: AI enables, the methodology governs, the researcher decides. In this framework, AI-generated suggestions are treated as research leads, never as evidence. No citation, transcription, abstraction, or proof statement enters the final argument until the human researcher has verified it against the underlying records. The tools expand what a single researcher can accomplish. The methodology ensures that expansion doesn’t come at the cost of evidentiary rigor. And the researcher — with their subject-matter expertise and professional judgment — remains the final authority on every conclusion.

    What It Produces

    I’ve applied this methodology across a demanding research project: tracing the English Tudor ancestry of a colonial American woman through nine armigerous lines converging in the clothier and gentry families of Cranbrook, Kent, c.1460–1620. The project has run across eighty-plus research sessions spanning multiple AI platforms, producing six publication-track articles across three peer-reviewed disciplines over the course of a year.

    The Courthope correction is perhaps the most telling example of the methodology in action. One AI platform generated a plausible claim about a generational relationship in a prominent Kent family pedigree. The adversarial auditor — a different platform with different training data — flagged it as suspicious. Investigation across platforms confirmed it was a false positive. But the investigation also uncovered something real: a generational error that had propagated through published sources for nearly two centuries, since William Courthope, Somerset Herald, first identified it — but the correction never reached the genealogical literature. The cross-platform workflow surfaced an error that had remained unresolved in the published line of transmission across multiple generations of scholarship. A correction article is in preparation for a leading county history journal.

    The methodology also produced a quantitative analysis of fiduciary trust networks among eleven Tudor clothier families — measuring whether formalized financial trust preceded or followed marriage alliances. Across twenty-one documented family pairings, formalized trust preceded marriage in every testable case, with zero counter-examples. An article is in preparation for a peer-reviewed social history journal.

    It identified a pedigree collapse — the same individual appearing twice in the ancestral tree through two different descent pathways — that connects nine documented armigerous lines through two intermarried gentry networks. A monograph documenting these lines with full GPS-compliant proof arguments is in preparation for a genealogical register.

    And it developed a diagnostic framework for assessing armigerous claims through extinct male lines — a methodological problem that traditional heraldic scholarship doesn’t address, but which cognatic-descent hereditary societies routinely encounter. An article is in preparation for a heraldic journal.

    The sycophancy researchers at MIT proposed mitigations that include “informing users of the possibility of model sycophancy.”[22] This methodology goes further. It doesn’t just warn researchers that AI might agree with them too readily — it builds structural disagreement into the workflow. The adversarial auditor’s job is literally to find reasons the analysis is wrong. When it can’t, that’s evidence of robustness. When it does, that’s the methodology working as designed.

    What’s Next

    A full article documenting this methodology — with detailed case studies, the complete GPS mapping, and the cross-platform verification protocol — is in preparation for submission to a peer-reviewed genealogical journal. This post describes the methodology and its general results; the specific evidence, proof arguments, and detailed analysis are reserved for peer-reviewed publication. The framework is portable: any field with an evidentiary standard — from legal scholarship to historical research to clinical case reporting — can use it to govern AI without tying the method to any one platform. The governance layer doesn’t depend on any specific AI platform — it governs whatever tools are in use, and it will continue to govern whatever tools replace them.

    I’ll be sharing more about specific findings — including the error that persisted across centuries of scholarship, the trust networks that preceded marriage, and the diagnostic framework for extinct male lines — in upcoming posts.

    Richard E. Rudd is an independent researcher. One of his current projects traces nine armigerous ancestral lines through two Tudor gentry networks in the Kentish Weald, c.1460–1620. This research was conducted using a multi-platform AI verification methodology; a full disclosure of AI tools and methods will appear in the published articles.


    Notes

    1. Salvagno, M., Taccone, F.S., & Gerli, A.G. (2023). “Artificial intelligence hallucinations.” Critical Care, 27(1):180. doi:10.1186/s13054-023-04473-y.
    2. Lee, J.H. & Shin, J. (2024). “How to Optimize Prompting for Large Language Models in Clinical Research.” Korean Journal of Radiology, 25(10):869–873. doi:10.3348/kjr.2024.0695. The authors document “prompt brittleness” — minor prompt variations producing substantially different outputs.
    3. Google (2025). Prompt Engineering White Paper. Available at kaggle.com/whitepaper-prompt-engineering. See also Anthropic’s published prompting documentation and OpenAI’s GPT Best Practices guide.
    4. Chandra, K., Kleiman-Weiner, M., Ragan-Kelley, J., & Tenenbaum, J.B. (2026). “Sycophantic Chatbots Cause Delusional Spiraling, Even in Ideal Bayesians.” (Preprint, arXiv 2602.19141.) MIT CSAIL, University of Washington, MIT Department of Brain & Cognitive Sciences.
    5. Hill, K. (2025). “They Asked ChatGPT Questions. The Answers Sent Them Spiraling.” The New York Times, 13 June 2025. The Human Line Project has documented nearly 300 cases of “AI psychosis” or “delusional spiraling.” Cited in Chandra et al. (2026).
    6. Chandra et al. (2026): “This effect persists in the face of two candidate mitigations: preventing chatbots from hallucinating false claims, and informing users of the possibility of model sycophancy.”
    7. Karpathy, A. LLM council concept. See e.g. https://x.com/karpathy — widely discussed in AI research community since 2023.
    8. Dr. Dominic Ng (@DrDominicNg), post on X, 17 March 2026, 1.8M views. https://x.com/DrDominicNg/status/2034252746996785213
    9. The 2005 PAL/CSS Freestyle Chess Tournament, in which amateurs using commodity hardware and chess engines defeated both grandmasters and dedicated supercomputers. Widely documented in chess history.
    10. Ng (2026). Current top engines (Stockfish, Elo ~3,653) exceed the highest human rating by approximately 800 points, making human intervention computationally harmful. See also Rao, V. (@VivekVRao1), X post, 18 March 2026.
    11. Board for Certification of Genealogists, Genealogy Standards (2d ed., 2019; rev. ed. 2021). The GPS comprises five interdependent elements. Element 4 (resolution of conflicts in evidence) and Element 5 (soundly reasoned conclusions) inherently require human judgment about ambiguous and contradictory evidence.
    12. ISA 610 governs the use of internal auditors’ work; SOX Section 404 requires management assessment of internal controls over financial reporting. The principle — that those who produce work should not be the sole evaluators of that work — provides the conceptual parallel for multi-platform AI verification.
    13. Board for Certification of Genealogists, Genealogy Standards (2d ed., 2019; rev. ed. 2021).
    14. Coalition for Responsible AI in Genealogy (CRAIGEN). Guiding principles available at craigen.org. CRAIGEN’s five principles address accuracy, disclosure, privacy, education, and compliance with existing genealogical standards.
    15. Little, S. NGS AI Program Director. Presentations include “Uses of AI in Genealogy” (RootsTech, NGS conferences, 2024–2026) and the “Genealogy Narrative Assistant” project for GPS-compliant AI-assisted writing. See aigenealogyinsights.com. Cross-platform GEDCOM comparison: Little, S. “Turn Anything into a GEDCOM File — With Any AI Tool.” Vibe Genealogy (Substack), 31 March 2026. vibegenealogy.ai/p/turn-anything-into-a-gedcom-file. Both platforms independently caught contradictory evidence in the source and documented the discrepancy rather than silently resolving it.
    16. Ancestors and Algorithms: AI for Genealogy podcast, hosted by Brian. Episode 30 (March 2026) maps GPS elements to specific AI platforms by functional strength. Available at ancestorsandai.com.
    17. See e.g. Shan et al., “Community-Driven AI Support for Genealogy Research” (Virginia Tech, CSCW 2023 workshop); Rice University Fondren Fellows project (2026), teaching cross-platform AI output verification. Multi-agent AI pipelines for fact verification: arXiv 2510.22751 (2025).
    18. Liang, P., et al. (2022). “Holistic Evaluation of Language Models (HELM).” Stanford CRFM, arXiv 2211.09110. LMSYS Chatbot Arena: Chiang, W.-L., Zheng, L., et al. (2024), arXiv 2403.04132. As of early 2026, no single model dominates across all task categories.
    19. Zhou, Y., et al. (2022). “Large Language Models Are Human-Level Prompt Engineers.” arXiv 2211.01910. Google DeepMind’s OPRO: Yang, C., et al. (2023). “Large Language Models as Optimizers.” arXiv 2309.03409. OPRO’s optimized instructions exceeded zero-shot human-crafted prompts by up to 8% on GSM8K, with substantially larger gains on Big-Bench Hard tasks.
    20. The MAPS framework (arXiv 2501.01329, 2025) demonstrates, in the context of software test generation, that prompts optimized for one LLM perform measurably differently on others — confirming that platform-specific optimization yields real advantages.
    21. The GPS requirement for “resolution of conflicts in evidence” (Element 4) inherently demands consultation of multiple independent sources. The methodology extends this principle from sources to tools.
    22. Chandra et al. (2026), abstract.