Opsgenie shuts down April 2027. Faultline replaces paging and adds the detection layer Opsgenie never had. Migration guide →
Skip to content
Roadmap

Where Faultline is headed.

Faultline is built one capability per release, with an AI-native bias: detect earlier, explain the failure, and close the loop. The track below runs solid through what's already live and turns to a dashed projection for what we're still exploring.

Last updated July 20, 2026 · Directional, not a contract. Priorities shift as we learn.

Available nowLive in every plan today.

Anomaly detection

Robust statistical baselines (median + MAD over same-hour-of-week history) flag latency and uptime drift before a fixed threshold ever trips, and open an advisory signal or, past a high-confidence gate, an incident.

Incident memory

Every resolved incident is embedded and stored. New diagnoses retrieve the most similar past incidents and cite them, so remediation gets smarter with every fix you close.

AI incident copilot

Incidents summarize themselves (what happened, blast radius, likely cause), draft a blameless post-mortem on resolve, and recommend a runbook, all from the dashboard. MTTA and MTTR stay computed in code; only the prose is drafted.

Graduated-autonomy remediation

Each runbook picks an autonomy level. Autonomous runs verify recovery afterward and page a human if the service didn't actually come back. Risky runbooks stay behind an explicit approval gate.

Alert correlation

When a shared dependency takes down many services at once, Faultline collapses the storm into a single page with correlated children instead of paging you N times at the worst possible moment.

Natural-language monitor creation

Describe what you want watched in plain English and Faultline parses it into a configured monitor with the right type, interval, and thresholds.

Agent-operable everywhere

One agent loop behind an MCP server, the faultline CLI, and a Slack bot. Point Claude at Faultline and it can inspect health, diagnose incidents, and run runbooks behind your approval.

In-app copilot console

The same agent loop that powers the CLI and Slack, streamed straight into the dashboard as a conversational console. Ask about any service or incident, watch it reason over live tools, and approve mutating actions inline.

Next upCommitted and near-term.

Multi-region quorum probes

Run checks from several regions and require a quorum before calling a service down, so a single bad vantage point can't page you. Faultline runs from one checker fleet today; multi-region quorum is next.

Synthetic customer-flow monitoring

Scripted multi-step journeys (login to cart to checkout) run on a schedule, so you catch a broken flow the way a customer would hit it, not just a single failing endpoint.

AI-drafted runbooks from history

Incident memory already knows how your recurring failures get fixed. Next we turn that into a ready-to-review runbook draft, so the third time a pattern shows up, the playbook is already written.

Service dependency graph

Auto-infer your topology from correlated failures and traffic, then use it to sharpen blast-radius estimates and route correlation toward the true root cause instead of the loudest symptom.

SLO and error-budget tracking

Define SLOs per service and track error-budget burn, so severity and alerting follow the budget you're actually spending instead of a flat threshold.

Predictive saturation forecasts

Project capacity and saturation trends forward so Faultline can warn you days ahead, 'connection-pool exhaustion in ~3 days at this rate,' rather than only paging once it breaks.

Monitors-as-code

Define monitors, thresholds, and runbooks in version-controlled config and apply them like infrastructure, with review, diffs, and rollback.

SOC 2 Type II

A formal security and compliance program, SOC 2 Type II, so Growth and Enterprise buyers have the security artifact their procurement asks for. Planned, not yet certified.

Self-hosted Enterprise deployment

Run Faultline inside your own VPC or VNET, for teams that can't send infrastructure signal to a managed cloud. Planned for the Enterprise tier.

ExploringIdeas we're prototyping. No dates yet.

Recurring-incident insights

Cluster the embedding store to surface 'this is the 4th time this quarter' patterns, with the common thread and the fixes that actually held.

Auto-narrated status updates

Draft customer-facing status-page posts from the internal incident, tone-controlled and approval-gated, so the public update keeps pace with the response.

On-call handoff briefings

An AI brief for whoever's next on call: open incidents, what changed this shift, and what to keep an eye on.

Adaptive thresholds

Learn which alerts get acted on versus silenced and tune per-service thresholds toward your real noise budget.

Multi-hypothesis investigation

Race several agent investigators against competing root-cause theories in parallel and surface the one the evidence supports.