Incident Response

Private Preview

Set up Droid to respond to production alerts in Slack and resolve or investigate incidents autonomously.

Note

Incident Response is in Private Preview. To request access, fill out the form at factory.ai/contact or contact your Factory account team.

Overview

When an alert lands in a configured Slack channel, an Incident Response automation starts a linked Droid session so Droid can investigate the incident with the right context, computer, and instructions.

Droid can work toward an RCA, open a PR for a fix when appropriate, and update an incident runbook so future investigations of the same alert type start faster.

Quickstart: set up from the Factory App

Incident Response is a guided setup for a Slack automation in the Factory App.

  1. 1
    Invite Factory to the incident channel

    In Slack, open the channel where alert bots post incidents and run:

    /invite @Factory

    If you have not connected Slack yet, set up the Slack integration first. Only channels where Factory has been invited appear in the channel picker.

  2. 2
    Open the Incident Response setup

    In the Factory App, open Automations from the sidebar and select New automation. Under Guided setups, select Incident Response.

    If New automation opens a menu instead, choose Create with Template, then select Incident Response under Slack.

  3. 3
    Name the automation and select trigger channels

    The name defaults to Incident Response. Change it if you plan to set up more than one. Under Trigger, search for and select the incident channel. Trigger on messages from defaults to Bots only, which matches alert bots posting incidents; change it only if humans should also trigger investigations.

  4. 4
    Choose where it runs

    Under Where it runs, choose the identity, computer, and session visibility for each investigation:

    • Run as service account? appears if you have the Manager or Owner role and your organization can use service accounts. A service account is recommended for incident response so the computer and session interactivity are shared across the team. Leave it at No selection (default) to run as yourself.
    • Run on sets the computer where Droid investigates. Pick the Droid Computer with the right repositories, observability tools, and anything else that can help Droid. With a service account selected, only its computers are listed.
    • Session privacy defaults to Team, so others in your organization can open the investigation sessions. Choose Private to limit them to the run identity.
  5. 5
    Configure MCP servers

    After you pick a remote computer, an MCP servers row appears. Select Configure to add and authenticate the MCP servers for the observability tools and context sources Droid needs to respond to incidents. MCP servers are configured per computer, and setup needs a live connection to it. If setup fails, retry or pick another computer.

  6. 6
    Review the prompt and model

    The template pre-fills a prompt that is added to every session a matching message starts:

    Use the `/incident` skill for Slack incidents and run an RCA if applicable.
    If the thread message is an incident or asking for RCA/root-cause analysis, invoke the `incident` skill before investigating.

    Keep the default or add channel-specific instructions. Optionally, pick a model and reasoning level from the controls under the prompt.

  7. 7
    Choose who can see it and create

    Use Share in the form header to keep the automation private or share it with your organization. Choosing a service account under Run as service account? shares it automatically. Review any Additional settings, then select Create.

  8. 8
    Send a test alert

    Post a test top-level alert from the same kind of bot that will send real incidents. Confirm the automation starts a Droid session and links it from the Slack thread. You can review runs in the automation's Events feed.

Configuration details

Customizing the prompt

The automation prompt is added to every session triggered from the channel. Keep it specific to incident work:

  • Tell Droid to run RCA and use the /incident workflow when appropriate.
  • Mention the primary services or repositories for that channel.
  • Point Droid at runbooks, dashboards, or common alert sources.
  • State escalation expectations, such as when to summarize uncertainty instead of taking action.

Avoid prompts that make the automation a general-purpose Slack surface. Incident Response works best when the prompt is narrow and operational. For broader Slack workflows, create a separate Slack automation instead.

Channel selection and trigger source

Incident Response is designed for channels where incident alerts arrive as top-level messages. Thread replies never start new sessions, which prevents ordinary discussion from repeatedly launching Droids.

For best results, use a dedicated channel such as #incidents, #alerts-production, or a service-specific incident channel. If your team opens a new public channel for each incident, enter a name pattern such as incident-* in the channel picker. Factory joins matching channels as they're created, so each new incident channel is covered automatically. See Slack channels and messages for pattern rules.

The channel picker only shows channels where Factory has been invited. Trigger on messages from defaults to Bots only so ordinary messages in the channel do not trigger investigations; widen it only if humans should also be able to start incident sessions.

To investigate only some alerts, open Additional settings. Keywords limits runs to alerts that mention one of your terms, such as a severity or service name, and Exclude keywords and Exclude senders skip noisy alerts. Keyword matching also reads the blocks and attachments that alert bots often post their content in.

Session privacy and model

Session privacy controls who can see triggered sessions. Incident Response defaults to Team, so other members can review sessions, which is useful for incident handoff and postmortem review. Private sessions are visible only to the run identity.

Use the default model unless your team has a known preference for incident analysis. For complex production incidents, choose a stronger reasoning model or a higher reasoning level.

Managing the automation

Open the automation from Automations to see its settings and its Events feed of runs and alerts. As the owner, you can select Edit to change its name, channels, trigger source, prompt, model, session privacy, and filters. The ⋯ menu can Pause, Resume, Share, or Delete it. Pause the automation instead of deleting it when you want to temporarily stop incident sessions.

The run identity and computer can't change after you create the automation. To move investigations to another computer or identity, create a new Incident Response automation, then delete the old one. For more options, see Manage an automation.

How the incident flow works

Once configured, Incident Response follows this flow:

  1. 1
    An alert bot posts a top-level message in the configured Slack channel.
  2. 2
    The automation starts a Droid session using its run identity, computer, session privacy, model, and prompt.
  3. 3
    Droid receives the Slack alert context and checks for any matching incident-guidelines runbook guidance.
  4. 4
    Droid investigates with the configured tools and should invoke the /incident workflow when RCA is appropriate.
  5. 5
    Factory posts status and session links back into Slack, and Droid can save reusable learnings for similar incidents.

The incident-guidelines runbook is stored locally at .factory/skills/incident-guidelines/SKILL.md for team reuse or ~/.factory/skills/incident-guidelines/SKILL.md for personal reuse. It should store reusable investigation guidance, not secrets or one-off RCA details.

Troubleshooting