# Incident Response

Set up Droid to respond to production alerts in Slack and resolve or investigate incidents autonomously.

<Note>
  Incident Response is in Private Preview. To request access, fill out the form at <a href="https://factory.com/contact">factory.ai/contact</a> or contact your Factory account team.
</Note>

## Overview

When an alert lands in a configured Slack channel, an Incident Response automation starts a linked Droid session so Droid can investigate the incident with the right context, computer, and instructions.

Droid can work toward an RCA, open a PR for a fix when appropriate, and update an incident runbook so future investigations of the same alert type start faster.

## Quickstart: set up from the Factory App

Incident Response is a guided setup for a Slack [automation](/software-factory/automations) in the Factory App.

<Steps>
  <Step title="Invite Factory to the incident channel">
    In Slack, open the channel where alert bots post incidents and run:

    ```text
    /invite @Factory
    ```

    If you have not connected Slack yet, set up the [Slack integration](/remote-delegations/slack) first. Only channels where Factory has been invited appear in the channel picker.
  </Step>

  <Step title="Open the Incident Response setup">
    In the Factory App, open <a href="https://app.factory.ai/automations"><strong>Automations</strong></a> from the sidebar and select **New automation**. Under **Guided setups**, select **Incident Response**.

    If **New automation** opens a menu instead, choose **Create with Template**, then select **Incident Response** under **Slack**.
  </Step>

  <Step title="Name the automation and select trigger channels">
    The name defaults to **Incident Response**. Change it if you plan to set up more than one. Under **Trigger**, search for and select the incident channel. **Trigger on messages from** defaults to **Bots only**, which matches alert bots posting incidents; change it only if humans should also trigger investigations.
  </Step>

  <Step title="Choose where it runs">
    Under **Where it runs**, choose the identity, computer, and session visibility for each investigation:

    - **Run as service account?** appears if you have the Manager or Owner role and your organization can use [service accounts](/enterprise/identity-and-access#service-accounts). A service account is recommended for incident response so the computer and session interactivity are shared across the team. Leave it at **No selection (default)** to run as yourself.
    - **Run on** sets the computer where Droid investigates. Pick the [Droid Computer](/droid-computers/overview) with the right repositories, observability tools, and anything else that can help Droid. With a service account selected, only its computers are listed.
    - **Session privacy** defaults to **Team**, so others in your organization can open the investigation sessions. Choose **Private** to limit them to the run identity.
  </Step>

  <Step title="Configure MCP servers">
    After you pick a remote computer, an **MCP servers** row appears. Select **Configure** to add and authenticate the MCP servers for the observability tools and context sources Droid needs to respond to incidents. MCP servers are configured per computer, and setup needs a live connection to it. If setup fails, retry or pick another computer.
  </Step>

  <Step title="Review the prompt and model">
    The template pre-fills a prompt that is added to every session a matching message starts:

    ```text
    Use the `/incident` skill for Slack incidents and run an RCA if applicable.
    If the thread message is an incident or asking for RCA/root-cause analysis, invoke the `incident` skill before investigating.
    ```

    Keep the default or add channel-specific instructions. Optionally, pick a model and reasoning level from the controls under the prompt.
  </Step>

  <Step title="Choose who can see it and create">
    Use **Share** in the form header to keep the automation private or share it with your organization. Choosing a service account under **Run as service account?** shares it automatically. Review any **Additional settings**, then select **Create**.
  </Step>

  <Step title="Send a test alert">
    Post a test top-level alert from the same kind of bot that will send real incidents. Confirm the automation starts a Droid session and links it from the Slack thread. You can review runs in the automation's **Events** feed.
  </Step>
</Steps>

## Configuration details

### Customizing the prompt

The automation prompt is added to every session triggered from the channel. Keep it specific to incident work:

- Tell Droid to run RCA and use the `/incident` workflow when appropriate.
- Mention the primary services or repositories for that channel.
- Point Droid at runbooks, dashboards, or common alert sources.
- State escalation expectations, such as when to summarize uncertainty instead of taking action.

Avoid prompts that make the automation a general-purpose Slack surface. Incident Response works best when the prompt is narrow and operational. For broader Slack workflows, create a separate [Slack automation](/software-factory/automations#start-from-scratch) instead.

### Channel selection and trigger source

Incident Response is designed for channels where incident alerts arrive as top-level messages. Thread replies never start new sessions, which prevents ordinary discussion from repeatedly launching Droids.

For best results, use a dedicated channel such as `#incidents`, `#alerts-production`, or a service-specific incident channel. If your team opens a new public channel for each incident, enter a name pattern such as `incident-*` in the channel picker. Factory joins matching channels as they're created, so each new incident channel is covered automatically. See [Slack channels and messages](/software-factory/automations#slack-channels-and-messages) for pattern rules.

The channel picker only shows channels where Factory has been invited. **Trigger on messages from** defaults to **Bots only** so ordinary messages in the channel do not trigger investigations; widen it only if humans should also be able to start incident sessions.

To investigate only some alerts, open **Additional settings**. **Keywords** limits runs to alerts that mention one of your terms, such as a severity or service name, and **Exclude keywords** and **Exclude senders** skip noisy alerts. Keyword matching also reads the blocks and attachments that alert bots often post their content in.

### Session privacy and model

**Session privacy** controls who can see triggered sessions. Incident Response defaults to **Team**, so other members can review sessions, which is useful for incident handoff and postmortem review. **Private** sessions are visible only to the run identity.

Use the default model unless your team has a known preference for incident analysis. For complex production incidents, choose a stronger reasoning model or a higher reasoning level.

### Managing the automation

Open the automation from <a href="https://app.factory.ai/automations"><strong>Automations</strong></a> to see its settings and its **Events** feed of runs and alerts. As the owner, you can select **Edit** to change its name, channels, trigger source, prompt, model, session privacy, and filters. The **⋯** menu can **Pause**, **Resume**, **Share**, or **Delete** it. Pause the automation instead of deleting it when you want to temporarily stop incident sessions.

The run identity and computer can't change after you create the automation. To move investigations to another computer or identity, create a new Incident Response automation, then delete the old one. For more options, see [Manage an automation](/software-factory/automations#manage-an-automation).

## How the incident flow works

Once configured, Incident Response follows this flow:

1. An alert bot posts a top-level message in the configured Slack channel.
2. The automation starts a Droid session using its run identity, computer, session privacy, model, and prompt.
3. Droid receives the Slack alert context and checks for any matching `incident-guidelines` runbook guidance.
4. Droid investigates with the configured tools and should invoke the `/incident` workflow when RCA is appropriate.
5. Factory posts status and session links back into Slack, and Droid can save reusable learnings for similar incidents.

The `incident-guidelines` runbook is stored locally at `.factory/skills/incident-guidelines/SKILL.md` for team reuse or `~/.factory/skills/incident-guidelines/SKILL.md` for personal reuse. It should store reusable investigation guidance, not secrets or one-off RCA details.

## Troubleshooting

<Troubleshooting>
  <TroubleshootingItem title="Slack channel does not appear in the channel picker">
    Invite Factory to the channel with `/invite @Factory`, then search for the channel again in the automation's trigger settings. Private channels must invite Factory explicitly.
  </TroubleshootingItem>

  <TroubleshootingItem title="Cannot create the automation">
    After you select **Create**, the form header explains what is missing. Make sure you selected at least one channel and kept a prompt. If **Run as service account?** is set, pick one of the service account's computers under **Run on** and wait for it to connect.
  </TroubleshootingItem>

  <TroubleshootingItem title="No session starts after a test message">
    Confirm the message is a top-level alert in the configured channel, not a thread reply, and that the sender matches the **Trigger on messages from** setting. Check that the alert passes any keyword and sender filters under **Additional settings**, and that the automation is not paused.
  </TroubleshootingItem>

  <TroubleshootingItem title="Factory replies that no computer is configured">
    The automation has no computer to run on. Create a new Incident Response automation with a **Run on** computer, then delete the old one. For a service-account automation, the service account must also be active and still own the selected computer.
  </TroubleshootingItem>

  <TroubleshootingItem title="Droid starts but cannot investigate the issue">
    Check that the selected computer has the right repository, tooling, and credentials. Update the automation prompt with links to runbooks or dashboards Droid should use first.
  </TroubleshootingItem>

  <TroubleshootingItem title="RCA is too generic">
    Add more channel-specific context to the prompt: service names, alert sources, dashboard links, log query examples, repository paths, and expected RCA format.
  </TroubleshootingItem>
</Troubleshooting>

<RelatedLinks>
  <RelatedLink href='/remote-delegations/slack' title='Slack'>
    Connect Slack and invite Factory to your incident channels.
  </RelatedLink>
  <RelatedLink href='/droid-computers/overview' title='Droid Computers'>
    Configure persistent environments for investigations.
  </RelatedLink>
  <RelatedLink href='/harness/skills' title='Skills'>
    Learn how reusable workflows like incident investigation guide Droid.
  </RelatedLink>
  <RelatedLink href='/autonomy-and-safety/auto-run' title='Autonomy Level'>
    Understand how Droid runs work without repeated approvals.
  </RelatedLink>
</RelatedLinks>
