Security Review

Run security-focused PR reviews and full-codebase audits with Droid using STRIDE, OWASP, and supply-chain methodology.

Droid security review is a dedicated security workflow for finding high-confidence vulnerabilities in pull requests or across an entire repository. It can run locally from the CLI or automatically in GitHub Actions.

PR security review

Review only the pull request diff, trace changed data flows, and post inline security findings with severity and suggested fixes.

Full-codebase audit

Audit every source file in the repository, group files for parallel review, and produce a structured report of validated findings.

Run a full-codebase audit

For the most thorough security results, run the audit inside a Mission. Missions plan the audit upfront, fan out work across orchestrated agents, and validate findings at each milestone, which produces dramatically deeper coverage than a single-session run.

From any Droid session, enter a mission and kick off the security review:

/missions
/security-review deep audit

Periodic scan in CI

Run the same mission-based audit on a schedule by invoking droid exec --mission from a workflow. The audit writes its full output under ~/security-audits/<slug>-<YYYYMMDD>/ on the runner, so add an actions/upload-artifact step to preserve findings after the runner exits:

YAML
on:
  schedule:
    - cron: '0 6 * * 1'
 
jobs:
  audit:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - name: Install Droid CLI
        run: |
          curl -fsSL https://app.factory.ai/cli | sh
          echo "$HOME/.local/bin" >> "$GITHUB_PATH"
      - name: Run deep security review
        env:
          FACTORY_API_KEY: ${{ secrets.FACTORY_API_KEY }}
        run: |
          droid exec --mission --auto high -m claude-opus-4-7 \
            "/security-review across the entire repository"
      - name: Upload security review output
        if: always()
        uses: actions/upload-artifact@v4
        with:
          name: deep-security-review-${{ github.run_id }}
          path: ~/security-audits/
          if-no-files-found: warn
          retention-days: 90

Run locally on a diff

To review the current diff in your working tree or branch from the CLI, run the built-in skill in any Droid session:

/security-review local diff

When invoked on a diff, Droid traces changed data flows across authentication, authorization, validation, database, network, filesystem, and LLM boundaries, and reports validated findings inline with severity and suggested fixes.

Run in GitHub CI on pull requests

With Droid Action, comment on a pull request to trigger an on-demand security review:

@droid security

To run security review automatically on every non-draft PR, add automatic_security_review: true to your review workflow:

YAML
- name: Run Droid Auto Review
  uses: Factory-AI/droid-action@main
  with:
    factory_api_key: ${{ secrets.FACTORY_API_KEY }}
    automatic_review: true
    automatic_security_review: true

When automatic_review and automatic_security_review are both enabled, Droid runs the security pass alongside the standard code review and includes the security summary in the PR feedback.

Configuration

These are the Droid Action security inputs currently wired for the workflows documented on this page:

InputDefaultDescription
automatic_security_reviewfalseRun security review automatically on PRs without requiring @droid security.
security_model""Advanced. Override the model used for security review candidate generation and full-repository scans. Falls back to review_model if unset; if both are empty, PR security reviews follow the same review_depth preset as code review. See Advanced: model overrides.

Severity threshold, blocking, team notification, and scheduled-scan inputs are declared in action.yml.

Custom guidelines

Security review applies repository-specific rules from a security-review-guidelines skill. This is a separate skill from the review-guidelines skill used by code review; each review type reads only its own guidelines.

Create .factory/skills/security-review-guidelines/SKILL.md in your repository:

Markdown
---
name: security-review-guidelines
description: Repository-specific security review rules for this codebase.
---
 
Additional security checks for this codebase:
 
- Every HTTP handler under `src/api/` must call `requireOrgMember()` before reading tenant data. Report any handler that does not as High.
- `packages/billing/` is in PCI scope. Report card data written to logs as Critical.
- `scripts/` is internal tooling that never receives untrusted input. Do not report command injection there.

How guidelines are applied

  1. 1
    When the security-review skill is invoked (locally, in a mission, or from Droid Action), the Droid CLI checks the skills available in the session for a valid, enabled skill named security-review-guidelines.
  2. 2
    If one exists, the CLI appends an instruction to the security review methodology telling Droid to invoke the guidelines skill and to have every subagent it spawns invoke it too. In a deep audit, this includes every lieutenant and jury pass.
  3. 3
    The guidelines take priority over the shared STRIDE, OWASP, and supply-chain methodology when they conflict. Use them to add checks, narrow or widen scope, suppress accepted patterns, or change how severity is assigned for your codebase.
  4. 4
    If no security-review-guidelines skill exists, or it is invalid or disabled, the methodology runs unchanged.

Because the CLI performs this step, guidelines are picked up everywhere the skill runs: /security-review in the CLI, @droid security and automatic_security_review in Droid Action, and scheduled scans. No workflow changes are needed.

What guidelines cannot change

Droid treats the guidelines skill as untrusted, repository-controlled input. Guidelines can refine what is reviewed and how findings are prioritized, but they cannot relax the skill's safety invariants. Droid will not follow a guideline that asks it to upload or transmit findings, write outside the local audit directory, run destructive or out-of-scope commands, or skip consent and proof-of-concept execution gates. A guideline that conflicts with these invariants is ignored and the invariant is kept.

Methodology

Security review uses the built-in security-review skill. In PR automation, Droid Action runs a dedicated security-reviewer subagent that loads this methodology before reading files, then traces changed data flows across authentication, authorization, validation, database, network, filesystem, and LLM boundaries.

The methodology applies multiple security frameworks together:

  • STRIDE threat modeling: Spoofing, Tampering, Repudiation, Information Disclosure, Denial of Service, and Elevation of Privilege.
  • OWASP Top 10:2021: Broken access control, cryptographic failures, injection, insecure design, misconfiguration, vulnerable components, authentication failures, integrity failures, logging failures, and SSRF.
  • OWASP Top 10 for LLM Applications:2025: prompt injection, sensitive information disclosure, insecure LLM output handling, excessive agency, vector/embedding weaknesses, and other AI-specific risks when the codebase uses LLMs.
  • Supply-chain analysis: dependency manifest and lockfile review, including typosquatting signals, install scripts, overly broad version ranges, and newly published packages.
  • Repository threat-model context: if .factory/threat-model.md exists, Droid uses it as the attack-surface map.

Review pipeline

Security review uses a two-pass workflow:

  1. 1
    Candidate generation: Droid reads the diff or codebase, identifies security-relevant areas, traces untrusted input across trust boundaries, and produces candidate vulnerabilities.
  2. 2
    Validation: Droid re-checks each candidate for reachability, exploitability, existing controls, and false positives before reporting it.

Findings are reported only when there is a realistic exploit path, such as an injection vulnerability, missing authentication or authorization on a sensitive operation, hardcoded secret, data exposure, unsafe LLM output handling, or risky supply-chain change.

Severity levels

SeverityPriorityExamples
CriticalP0RCE, hardcoded production secret, auth bypass, unauthenticated admin endpoint
HighP1SQL injection behind auth, stored XSS, sensitive-data IDOR, very new dependency
MediumP2CSRF on state-changing operations, information disclosure, prompt injection behind auth
LowP3Minor security hardening with a concrete but low-impact exploit path