Skip to main content
Gradient Perspective
How do you backup an enterprise run by AI agents?
July 21, 2026

Enterprises will run thousands of AI agents. They’re going to need an ‘undo’ button.

Veeam, Commvault, Cohesity, Rubrik, Veritas and before them, HPE, Dell, IBM, and others spent 45+ years building backup and disaster recovery products to help businesses recover if they’re hit with ransomware, their data is corrupted, or they’re audited by a federal agency. These products generate $30B+ in revenue annually (and are large offerings for the hyperscalers as well).

These products take snapshots of a business’s data, replicate it across hard drives, and store the data in cloud (e.g., some enterprises still use tape storage!) or hot (e.g., Pure Storage’s flash storage recovery solutions) storage depending on how likely the data is to be needed again.

But those products were designed for a human-oriented world, where humans make decisions and update data manually, or at most, in large batches. The AI-native enterprise is going to run millions of agents simultaneously. And those agents are increasingly running read, write, update, and delete operations against core systems of record and making independent decisions based on context and LLM logic. Even the agents themselves change more rapidly (model post-training, harness optimization, context retrieval) than an individual employee as well.

The AI native business needs a new backup and disaster recovery solution.

$30B in legacy backup spend

Backup and disaster recovery are two overlapping markets: a fragmented ~$13B data-protection software market (backup) and a faster-growing ~$16B Disaster-recovery-as-a-service market (DRaaS). Together they make up a $30B+ industry of storing old data.

In fact, much of the 250,000+ exabytes stored today are sitting in backup & DSaaS repositories (IDC Global DataSpheres Report), and that number is now growing faster than ever thanks to genAI apps that capture & store more data (e.g., doctors visit transcriptions) that need to be backed up.

  • Backup is insurance against forgetting. Cold, cheap, long-lived data copies kept because a regulator requires it for security, auditability, legal support. They are read approximately never, but are often held for 5+ years.

  • Disaster recovery is insurance against stopping. Hot, recent copies whose job is to get a business running again after ransomware, a bad deploy, or a region falling over. Measured by how fast the business recovers (RTO - Recovery Time Objective) and how little the business loses (RPO - Recovery Point Objective). ‘DRaaS’, the most common type of disaster recovery solution (and fastest growing), is just paying someone else to sweat that hot path.

Why does anyone care? Because downtime is expensive. Over 90% of large enterprises say an hour of downtime costs $300K+, and 40% say it's north of $1M (ITIC 2024 Hourly Cost of Downtime Survey). The average ransomware incident drags on ~24 days (Coreware, via Statista), and IBM’s Cost of a Data Breach report pegs the average breach at $10M in the US.

Legacy backup solutions don’t work for AI-native enterprises

An enterprise run mostly by agents is at huge risk of going down if critical systems are compromised and existing backup and disaster recovery tools aren’t well placed to solve them.

Backup more information:

  • Accountability has to live inside the backup. When many agents, not a single person, make changes that corrupt data, "restore to yesterday" isn't enough. IT admins to know which agent did what, when, and why. Once a breach / issue happens that requires a reset, finding a clean backup that predates it becomes a challenge. Each DSaaS provider has their own IP here (algorithms combining snapshots for a healthy recovery point), and it only gets harder when agents corrupt the state continuously and the window between "last known good" and "already poisoned" collapses to minutes. Not to mention, there needs to be a tight link between any restore that happens and a resulting change to agent behavior / actions to ensure the data is corrupted again in minutes.

  • Constantly snapshotting agent state is not realistic. Memory and context change every few seconds, so backing it all up constantly would be too complex and costly. But there are moments worth capturing, especially in regulated industries where why an agent made a decision matters to an auditor. The challenge becomes deciding when a state is worth keeping.

Access backups for disaster recovery data more frequently and quickly:

  • Mission-critical agents need to recover fast. A stalled batch job waits until Monday, but an agent mid-task in a live workflow cannot, particularly if it is running a critical systems. That drags recovery onto the hot path, which may necessitate solutions like flash storage disaster recovery systems.

  • Long-running agents have complex dependencies. How does an agent that's accumulated valuable context (and burned real compute) restore itself when us-east-1 goes down and it loses everything?

The unit of protection is shifting from data at rest to an agent mid-task. That single change breaks most of the assumptions the incumbents were built on.

How infra providers are re-thinking backup and disaster response

Startups are moving backup and historical data tracking into the agent infrastructure itself, while legacy providers are rethinking how they expose stored data to AI agents.

  • Startups are building structure into the agent itself. Some version and branch each layer (Mesa for the filesystem, Ardent for the database, GitButler for code, Letta for memory) so recovery becomes an instant "undo," while Zep treats agent memory as a governed record with retention, legal hold, and audit built in.

  • Incumbent backup providers are retrofitting what they already sell. Eon and Cohesity's Gaia turn cold backups into a queryable data lake, while Veeam, Rubrik, Druva, Pure Storage, and NetApp bolt "detect and undo" and clean-room recovery onto existing snapshot engines. It's the fastest path to adoption, but mostly aimed at the reversible half of the problem.

The white space we’re excited about

We’re still in the early days of deploying AI agents at-scale throughout an enterprise but there are a few potential solutions to some of the emerging challenges for long-running, context-heavy agents that touch critical systems.

Build rewind and replay into the agent infrastructure itself to enable agent disaster recovery. Replace the traditional snapshot with a causally ordered log of every state change and tool call, plus deterministic replay so agents can fork from any point in time. Snapshots assume state lives in one place and changes slowly, but an agent's state is spread across memory, a vector store, a live plan, and a stream of tool calls, so a nightly snapshot just restores a lobotomized agent.

Tie backup state capture to external impact, and provide an adaptive policy engine (an infra, compliance, and regulation problem). Data at rest is reversible, but external actions are not. Databases can be restored, but agents can't un-send an email or un-charge a credit card. This presents challenges for the agent itself, but is actually a strong mechanism for determining when and if to snapshot agent state and surrounding systems. Agents should rarely need to snapshot state at all unless the agent(s) has actually made major changes to an outside system, so capture could be triggered by external impact rather than run on a fixed clock. That keeps the volume manageable and puts the backup emphasis on the inflection points that matter and agents that matter. The policy engine that ultimately drives this decisioning could be a significant source of stickiness and IP.

There is also a real opportunity to work with the regulatory ecosystem, including HHS OCR & NIST for healthcare and the SEC, FINRA, FFIEC, and CFTC for financial services, to define what backup and audit requirements should look like to maintain accountability in the age of AI, which is both a moat and a way to shape the category.

Enterprise Resilience is an important investment theme for the future

Ten years ago, one of the most exciting and lucrative investment themes was “Future of Work”, where we re-thought the tooling for knowledge workers with limitless access to SaaS tools delivered over the internet. Most folks will think of Zoom and Slack for collaboration, or Notion for productivity. But equally important was Okta, which brought MFA and SSO en masse to enterprises so that they could securely provide employees with access to those tools.

In this second phase of Future of Work that we are entering now, centered around human-to-agent and agent-to-agent collaboration, we once again need new enterprise resilience and security tooling to help enterprises adopt agents without significant risk to their business operations. Much of the focus has gone to new agent identity tooling, but we expect to see systems like Backup & disaster recovery become equally as critical to enterprise IT teams.

If you’re building in this space, reach out.