A code reviewer that reads before it comments.

ocra splits a change into focused review tasks. Each agent can only read your repository, has to quote the code it means, and is asked to say what it checked. Most runs end with a finding or two, many with none, and that is fine.

Node 22+, any model OpenCode supports

ocra review — example output
$ ocra review --from main
[ocra] Reviewing: Changes from main to HEAD
[ocra] 4 file(s) selected, 1 excluded · risk tier: full
[ocra] 2 bundle(s) (grouped)
[ocra] 4 review task(s), 2 reviewer/bundle pair(s) skipped
⋮
[ocra] security-1 completed in 63.2s · 2 finding(s)
[ocra] correctness-1 completed in 71.9s · 2 finding(s)
[ocra] Verified 3 finding(s), dropped 1 that the code disproves
⋮
Verdict: significant concerns
src/auth/session.ts
critical L42 Every session is treated as expired [verified] #b7d6c863
isExpired() compares expiresAt, stored in seconds, with Date.now() in milliseconds, so it is always true: loadSession() deletes every session and users are logged out right after signing in.
Suggestion: return session.expiresAt * 1000 < Date.now();
1 finding(s) (1 critical, 0 warning, 0 suggestion) · tokens: 315936 in (209152 cached), 9021 out, 8489 reasoning · $0.5134
01

What happens when you run ocra review

The steps that must not go wrong are plain, tested code. Models are only asked for judgment.

codemodel

  1. [00]

    Only the files worth reading

    selecttriageexample run
    • src/auth/session.tsread
    • src/auth/token.tsread
    • src/api/login.tsread
    • docs/auth.mdread
    • package-lock.jsonset aside: generated
    tiertriviallitefulltouches auth/

    Binaries, lock files, vendored and generated code, media, likely secrets and oversized diffs are set aside, each with a recorded reason. Sensitive paths such as auth/ or CI workflows, or more than 20 files, make the change full tier; otherwise its size decides between trivial, lite and full.

  2. [01]

    One task per group and reviewer

    bundlematrixexample run
    auth code3 filesdocs1 filecorrectnesssecurityskipperformanceskip
    4 review tasks2 skipped: no matching files

    From four files up, a light model groups files that belong together; smaller changes are split in code. Correctness always runs; security and performance join from the lite tier and skip docs and tests. Skipped pairs are recorded with their reason, and --plan lists every task before a model is paid.

  3. [02]

    Findings quote the code they mean

    reviewanchorexample run
    agent · correctness
    • read_fileon
    • read_diffon
    • code_searchon
    • editoff
    • bashoff
    • webfetchoff
    quotereturn session.expiresAt < Date.now();matched in the diffsrc/auth/session.ts:42

    An isolated agent per group and reviewer reads the exact revision under review with read-only tools, stops after 20 steps, and reports each issue by quoting code. ocra looks the quote up in the diff, then in the whole file, and pins it to those lines; a quote it cannot find stays on the file. The model never picks the line.

  4. [03]

    Only what holds up on a second read

    filterverifyjudgeverdictexample run
    • Every session is treated as expiredconfirmedcritical · session.ts:42
    • Expiry compares seconds with millisecondsmerged into #1
    • Token stays valid after logoutdisproved
    • user may be undefinedaccepted in memory
    verdict: significant concerns1 verified critical

    Findings the team accepted in ocra's memory are dropped before a model is paid to check them. A verifier rereads each one next to the code and drops what the code disproves; a judge merges the same root cause and drops nitpicks, but can never drop or downgrade a confirmed critical. The verdict is computed in code: only a critical the verifier confirmed blocks.

An example run, sped up. The pipeline, the output and the verdict are ocra's own; the model's answers were scripted for the recording.
02

Decisions we made on purpose

Each of these costs something. We think the trade is worth it.

  1. 01

    The model never picks the line.

    Models are bad at line numbers and good at quoting. So they quote, and ocra finds the lines.

  2. 02

    Reviewers are told what to leave alone.

    Style, speculation and unchanged code are out of scope for every reviewer, and each has its own list on top, such as missing tests for correctness. Reporting nothing is a valid outcome.

  3. 03

    No agent can write.

    Agents can read files, read diffs and search the repository. Editing, shell and web tools are switched off, along with every other OpenCode built-in.

  4. 04

    Your own setup stays out of the review.

    Your global OpenCode config, installed skills and instruction files are switched off before the first request, and the runtime gets only the environment variables it needs.

  5. 05

    Every attempt shows its bill.

    Steps, tool calls, tokens and cost for every review attempt as it finishes; the run's total adds cached tokens and every helper call.

  6. 06

    When a model falls over, the next one takes the task.

    Give each tier a list of models. Overloads move on to the next model, a model that keeps failing is paused for a while, short rate limits are waited out, and a model out of quota is dropped for the rest of the run.

03

What a finding looks like

An inline comment on the pull request, with enough to decide in a few seconds whether to fix it or dismiss it. What happens next is up to the code and the reviewers.

src/auth/session.tsExample pull request
41 export function isExpired(session: Session): boolean {
42+ return session.expiresAt < Date.now();1
43 }

github-actionsbot

Every session is treated as expiredcriticalverifiedcorrectness2

isExpired() compares expiresAt, stored in seconds, with Date.now() in milliseconds, so it is always true: loadSession() deletes every session and users are logged out right after signing in.3

Suggestion: return session.expiresAt * 1000 < Date.now();4

Example pull request

What a finding looks like

The finding stays open on the next push as long as its code is unchanged, even if no reviewer reports it again, so the same code keeps the same verdict.

  1. 1The line the agent quoted, found in the diff by ocra
  2. 2Severity, whether the verifier confirmed it, and which reviewer found it
  3. 3Why it is wrong, in plain words
  4. 4The smallest fix, when there is one
04

Your team's rules, as a plugin

The Git and GitHub adapters, the OpenCode runtime and the three reviewers that ship today are plugins too. Yours get the same small contract: register rules, reviewers, tools or listeners, and receive your own settings. Plugins run code, so they load only in local reviews; for pull requests, the same rules go in .ocra/rules.json on the base branch.

  • Three lifecycle hooks, run in a fixed order
  • Settings validated per plugin
  • Clashing or late registrations fail with the plugin's name
Read the plugin guide
tools/ocra-team-rules.mjs
export default {
  name: "team-rules",
  configure(ctx) {
    ctx.registerRules([{
      path: "services/**",
      rule: `${ctx.settings.team}: require idempotency keys`,
    }]);
  },
};
05

Where it stands

  1. M1Local reviewbuilt

    CLI, selection, bundling, anchoring, the correctness reviewer, OpenCode runtime, plugins, benchmark harness.

  2. M2More reviewersbuilt

    Security and performance reviewers, risk tiers, a review matrix, verification, a judge and a fixed verdict rubric.

  3. M3GitHubbuilt

    ocra review --pr and a GitHub Action: inline comments, one summary comment, and re-reviews of only what changed that resolve fixed threads and respect dismissals. Tested against a simulated GitHub API; the first live pull request comes with M7.

  4. M4Hardeningbuilt

    Circuit breakers per model, shared configuration over https, a review memory, and --ultra for recall.

  5. M5Measurebuilt

    A golden set of expected findings plus AACR-Bench. The first results are published with their limits: a small sample on one model family, with recall the weak point.

    Read the results
  6. M6Recallpaused

    Reviewers report every defect they can back with evidence, and verification and the judge own precision, one measured prompt change at a time. Paused until there is model credit to measure each change.

  7. M7Ship v0.1built

    v0.1 went to npm with provenance and installs with one line. The Action reviews real pull requests on two repositories.

  8. M8Untrusted pull requestsbuilt

    A threat model, a gated workflow for pull requests from forks, and an adversarial test set that plants instructions, links and commands in the change under review.

    Read the threat model
  9. M9Reach and trustnow

    GitLab merge requests, on GitLab.com or self-managed; SARIF for code scanning; a container image; your own OpenAI-compatible model endpoint; a security policy. Released in 0.2.0; next, a check on a live GitLab instance.

06

Try it on a repository you know

It runs against any Git repository on your machine, or on pull requests through the GitHub Action. Your model provider sees the change with its title and description, your AGENTS.md and review rules, and what the agents open or search in the repository; never files that look like secrets, or anything outside the repository.

Try it on a repository you know
zsh
npm install -g @open-cr-agent/cli
export GEMINI_API_KEY=...
export OCRA_MODEL_TOP=google/gemini-3.1-pro-preview
export OCRA_MODEL_STANDARD=google/gemini-3.5-flash
export OCRA_MODEL_LIGHT=google/gemini-flash-lite-latest
cd your-repository && ocra review --from main