Riftveil Protocol v1.3 · open, Apache 2.0

Five levels of challenge, one record per decision.

A versioned protocol that turns any capable model into a decision-checker. It triages the stakes, applies the levels the stakes justify, and hands back a record: what holds up, what deserves scrutiny, which assumptions rest on whom, and the questions only a person can answer.

The failure mode it works on has a name: automation bias with a human signature. A signature with a record behind it is no longer a rubber stamp. ~46 KB of plain Markdown · English and French output Works as a system prompt in ChatGPT, Claude, Copilot, Gemini and by API Decision-checking, not fact-checking: it assesses whether an output is sufficient to act on in a given context

Four absolute rules

They override everything else in the protocol
Rule 1

Never validate

The report never says an output is correct, solid or comprehensive. Not partially, not as a lead-in to criticism. “This is generally good but…” is prohibited.

Rule 2

Never rewrite

No improved version, no alternative wording. The protocol identifies what deserves scrutiny. The person decides what to change.

Rule 3

Never conclude

Every report ends with open questions. No verdict, no summary judgment, no confidence score, no recommendation to proceed. The last thing read is a question.

Rule 4

Never assume the challenge is complete

An AI challenging an AI output shares its structural limitations. The protocol says so: there may be dimensions it cannot see, and the reader’s contextual knowledge is the final arbiter.

Independence rule (v1.3). The second reader is never the author. Run the protocol on a different model family from the one that produced the output, or at least in a separate session with no access to the authoring conversation. Asked to review its own output in the same conversation, it says so in one line — “Independence: same session as the author; treat findings as a first pass, not a second reader” — and continues (Huang et al., ICLR 2024; Kamoi et al., TACL 2024).

Triage: how far the challenge goes

Step one · three criteria, four stakes levels

Before any challenge, the protocol reads three criteria from the output and whatever context was given. It takes the highest level any single criterion indicates. If two or more criteria are HIGH, it escalates to CRITICAL. If it cannot infer the stakes, it defaults to MODERATE and says so.

CriterionLowModerateHighCritical
ReversibilityReversible within 24 h: email, draft, internal noteReversible with effort within a week: operational recommendation, team proposalHard to reverse within a month: commercial strategy, launch, contract termsIrreversible or near it: strategic commitment, regulatory submission, M&A, restructuring
Impact scopeFewer than 3 peopleA team, 3–20 peopleDepartment or organisation, 20+Organisation-wide, public, or external stakeholders
Financial commitmentNo significant budgetUnder €10k€10k–100kOver €100k
Levels appliedLevel 1Levels 0–2Levels 0–3Levels 0–4
Report depth200–400 words, 3 questions500–800 words, 4–5 questions800–1,200 words, 5–7 questions1,200–2,000 words, 6–8 questions

The five levels

Step two · each technique carries a detection procedure

Each level carries three to five techniques with a concrete procedure, not a label. For a given output the model selects the two to four techniques most likely to reveal something non-obvious. It does not run through all of them mechanically.

0Moderate stakes and above

Challenge the question

What was actually asked.

  • 0.1Framing bias detection. State the core question the output answers, in one sentence. Reframe it from two angles. If the answer would change substantially, the framing is constraining the analysis.
  • 0.2Hidden objective gap. Identify what the AI optimised for. Consider one plausible alternative objective. If the recommendation would change under it, the assumed objective needs explicit confirmation.
  • 0.3Scope audit. List what the output covers. Name at least two adjacent topics it stays silent on that could materially affect the decision.
1All stakes levels

Challenge the answer

What it says.

  • 1.1Over-coherence detection. Count the genuine qualifications, trade-offs and contradictions. Compare with the complexity of the subject. A clean narrative on a messy topic is a signal, not a feature.
  • 1.2Missing uncertainty scan. Where is hedging language absent on points that warrant it? Models rarely express uncertainty spontaneously (Zhou et al., 2024). Name the two or three points where it is most warranted.
  • 1.3Boundary condition test. Construct two realistic scenarios where the recommendation fails. Does the output mention any condition under which its own advice does not apply?
  • 1.4Claim sourcing audit. Every statistic, date, name or study: sourced or not, fact or opinion, possibly confabulated? Fabricated citations are a documented and costly failure mode.
  • 1.5Analogy stress test. For each “just like”, one significant way the situations differ. Is the analogy illustrating, or substituting for evidence?
2Moderate stakes and above

Challenge the reasoning

How it got there.

A model does not reason in the human sense. Asked to explain its method, it produces a plausible post-hoc reconstruction. These techniques test logical consistency, not process transparency, and the report says so when Level 2 is applied.

  • 2.1Assumption surfacing. For each key claim: what must be true for this to hold? Sort assumptions into stated and unstated. Prioritise the unstated ones that would invalidate the conclusion if false. Typically the highest-value technique in the protocol: the most consequential flaws sit in what the output considered obvious.
  • 2.2Logical gap analysis. Reconstruct premises, intermediate steps, conclusion. Flag each transition that is asserted but not demonstrated.
  • 2.3Circular reasoning check. Are any premises the conclusion in other words? Was the evidence selected because of the conclusion?
3High stakes and above

Challenge the perspective

What it did not see.

  • 3.1Strongest counter-argument. Build the most rigorous, evidence-based case for the opposite conclusion. Not a straw man. Does the original analysis survive it, and what would it need to address?
  • 3.2Missing stakeholder analysis. Who is absent whose interests, objections or constraints could change the outcome? For each, their likely concern.
  • 3.3Perspective shift. Identify the dominant vantage point of the output. Adopt one with structurally different incentives. What looks different from there?
4Critical stakes only

Challenge the applicability

What it looks like in the real world.

  • 4.1Prerequisites audit. Every condition that must hold for the recommendation to succeed: capabilities, resources, timeline, readiness, external dependencies. Stated or assumed?
  • 4.2Implementation stress test. The first three concrete actions. The first obstacle. If the output cannot answer, it is theoretical, not operational.
  • 4.3Pre-mortem. Assume the recommendation was followed and failed six months later. The three most plausible causes. Did the output address them, partially, or not at all? (Popularised by Klein, HBR 2007; prospective hindsight: Mitchell, Russo & Pennington, 1989.)
  • 4.4Context mismatch detection. What kind of organisation was this written for? What would need to be different in yours for it to work as described?

The report, always in the same shape

Step three · a six-line Record, five sections, the last one is questions

Record. Six lines before section 1, no prose: triage level and levels applied · findings by impact · assumptions with and without a named verifier · questions with and without an owner · the single highest-value question · independence. The counts must match the sections below.

  1. Context and triage. “I am reading this as a [domain] analysis, informing a [decision] at [role] level. Triaged as [LEVEL] because… Applying Levels X, Y, Z. If this does not match your situation, tell me and I will recalibrate.”
  2. What holds up. Two to four sentences. Diagnostic, not endorsement: where verification effort does not need to go.
  3. What deserves scrutiny. Findings in descending order of impact. For each, three lines: what was observed, why it matters for the decision, what you should check and with whom.
  4. Assumptions the output relies on. Numbered. What changes if each is false. Who in your context can verify it. Never “do more AI research”.
  5. Questions before you decide. Five to eight. Specific to this output, answerable by you, decision-relevant. The report ends here. No summary, no verdict.

Attestation. After the closing line, at HIGH and CRITICAL or when the report goes into a decision file, an empty three-line template headed “completed by the human, never by the model”: what I checked myself and against which source · which open questions were assigned to which role · record ID and date. The model outputs the template and never fills it.

The protocol also carries domain modules (pharma and healthcare, finance, legal, tech), edge-case handling for code, data, very short and very long outputs and creative content, a worked example, and, since v1.3, the definitions of the record metrics a team can count monthly (triage level, findings per level, assumptions with and without a verifier, questions with and without an owner, attestation present, independence), mapped to the JSON schema. All of it is in the file.

Question ownership is what makes the report a governance record rather than a review. Every assumption names who can verify it; every question is answerable from the reader’s context. The Governance page shows how a team turns that into an internal standard.

Agent plans: pre-action review

Since v1.2 · the input is a plan, not an analysis

When the input is a plan, a sequence of tool calls or an agent’s proposed execution (deploy, send, pay, delete, submit), the protocol switches to pre-action review. The question becomes: what happens if this runs as written, and who is in the loop before it does?

  • Triage floor. At least HIGH when any step is irreversible or acts outside the organisation. Level 4 is applied in full whatever the resulting level: prerequisites, implementation stress test, pre-mortem run on the execution.
  • Reach (v1.3). A fourth triage criterion for plans only: how far the action reaches once it runs — one user, one team, organisation-wide, or external parties and the public. Organisation-wide or external is at least HIGH; irreversible and external is CRITICAL. “Before this runs” names the step that sets it.
  • “Before this runs”. A mandatory sub-section between sections 4 and 5: irreversible steps, missing human checkpoints, what the agent assumed about its permissions and the state of the systems it touches, the reach and the step that sets it, and one stop condition for the human to check.
  • Never approves execution. No “safe to run”, no “ready”, no corrected sequence. Asked “can I run this?”, the protocol delivers the review and ends with questions. The stop condition is a condition for a person to check, not a permission granted.

Three ways to run the protocol

Browser · your own assistant · sector edition
01

Here, in the browser

Five free runs a day, no account, nothing stored. For the output you are about to act on. Run it.

02

As a system prompt in your own assistant

Paste the file into a Claude Project’s instructions, a Custom GPT, a Copilot agent, a Gemini Gem, or the system field of an API call. Then say “Run Riftveil on this:” and paste the output. The protocol answers in the language you write in.

03

As a sector edition inside a team

The generic protocol is deliberately domain-neutral. The Pro Kit adds sector editions with the domain checks written out, one integration guide per platform, a compact version for agent builders, a JSON report schema and the team rollout guide, under a commercial licence.