Consistent

    Your documents aren't

    AiirGap reads your policies, SOPs, specs and controls and finds the statements that can't all be true

    A project in contradiction detection, run on one workstation

    Where search finds similar text, AiirGap classifies whether two statements can both be true and verifies every evidence quote against the source character by character.

    It Shows Its Work

    Source-verified evidence and per-token confidence for every finding

    Air-Gapped by Design

    Fully offline on one workstation. Your documents never touch the internet

    Honest About Uncertainty

    It says "look at this" rather than "this is broken" and filters false positives first

    How It Works

    Five stages turn a document set into verified, explainable conflict findings

    01

    Extract

    Pulls clean text from your documents; a vision model reads charts, scans and tables

    02

    Pair

    Finds passages likely to disagree, by meaning and by structure

    03

    Classify

    Decides whether both statements can be true, with confidence from the model's own logits

    04

    Verify

    A second pass filters false positives and diagnoses the nature of each conflict

    05

    Explain

    Source-verified evidence quotes, per-token confidence and a full audit trail

    Domain-agnostic by design:trained across 10 audited industry domains, from pharma and aviation to maritime and nuclear energy

    Contradiction is not one phenomenon. Statements can clash in what they measure, mandate, mean and entail.

    Each clash has its own logic. AiirGap brings a specialist reasoner to every fundamental form of disagreement.

    See What the Model Was Thinking

    Every verdict comes with a certainty card that shows how it was reached, so the reviewer sees the model's internals as well as its answer

    Verified Evidence Badges

    Quotes checked against the source character by character and badged verified, paraphrased or unverified

    Per-Token Confidence Sparklines

    See where the model was certain and where it guessed, token by token

    The SAE Radar

    Six calibrated axes of the model's internal state, read through the Jacobian lens for every finding

    The Verdict-Formation Trendline

    The Jacobian lens shows where in the network the verdict crystallized; verdicts that crystallized late get flagged for human attention

    Verdict certainty
    Retentionofoneyearcannotsatisfya90-daycap

    Per-token confidence (hover the words)

    Six calibrated internal axes

    L19layer 132

    Where the verdict crystallized

    "The model settled this verdict at layer 19 and never wavered."

    Built to be Audited

    Air-gapped by design, not by permission slip. Built for the questions a security review would ask

    Sealed Offline Bundle

    Installs from a sealed 16 GB bundle and makes zero external calls, verified

    Login Built In

    Always-on login, per-project access control, TOTP two-factor and SSO to your identity provider

    Neutrality, Measured

    Political neutrality is measured with a standing eval that every model swap must pass

    Full Audit Trail

    Every prompt, response and logprob stored; every verdict attributable, model or human

    Runs in Two Places

    An air-gapped workstation on Apple Silicon or your own GPU server on CUDA, from one codebase

    No Data Center Required

    The whole pipeline draws 23 W at the desk: 3.3 Wh per 1,000 comparisons, measured, with no token meter and no cloud

    Measured, Not Marketed

    Every number below traces to a reproducible engineering record. On our held-out evaluation, AiirGap found every planted contradiction and raised zero false alarms.

    The entire EU AI Act for a third of a phone charge

    459
    pages read
    12
    minutes, end to end
    5
    watt-hours

    One MacBook Pro, 740 comparisons, five contradictions flagged, measured at match sensitivity 0.93. Among them: a deadline that runs four years in one article and seven in another. Each verdict used 37× less energy than a single chatbot prompt, by Google's and OpenAI's own numbers.

    ~1s

    Per verdict, everything on

    0.958s per comparison through the full pipeline, with quote verification, model internals and the Jacobian lens live

    ~34%

    Of our training data rejected

    An adversarial label audit condemned a third of our hand-authored pairs across 10 domains

    Chatbot energy per Google (0.24 Wh, median Gemini text prompt) and OpenAI (0.34 Wh, average ChatGPT query); a full AiirGap verdict measured 6.5 mWh. Phone charge relative to the iPhone 16's 13.8 Wh battery. Run measured at match sensitivity 0.93; the default of 0.90 pairs roughly twice as many comparisons and takes proportionally longer. Flagged findings are review queues for humans, not adjudicated errors in the Act.

    Follow the project

    The next experiment puts two sources that disagree in front of the model and reads the odds. Leave an email address and hear how it comes out.