← All dispatches
Dispatches · #intelligence · ScrapeOps

DataBahn's $40M Round Shows What AI Data Pipelines Cost

August 1, 2026 · Abhishek Gupta
Infographic: DataBahn raised $40M Series B on 600+ data sources routed, 400% year-over-year revenue growth, and 0% customer churn for AI data pipelines

Insight Partners just put $40 million into a company whose entire pitch is that most of the data flowing into your AI models shouldn't be there. DataBahn closed a Series B on July 30, 2026, taking its total raised to $59 million. The bet: AI data pipelines are now expensive enough that filtering what reaches a model is a standalone business.

The short version

  • DataBahn raised a $40M Series B led by Insight Partners on July 30, 2026, bringing total funding to $59M since its 2024 founding.
  • The platform ingests and normalizes telemetry from 600+ enterprise sources before any of it reaches an AI application or model.
  • Customers include MVB Bank and the Canada Pension Plan Investment Board; revenue grew 400%+ year over year with 0% churn and a 97% proof-of-concept win rate.
  • The pitch isn't "store more data" — it's "route less of it," cutting cloud egress, storage, and inference bills.
  • It's the same problem ScrapeOps solves from the opposite end: the messy public web instead of messy internal systems.

Why is a data-routing company worth $40M right now?

Because the cost of feeding an AI model badly has become visible on a P&L. Enterprises spent 2024–2025 dumping every log, event, and telemetry stream into a data lake on the theory that more context makes AI smarter.

That theory has a bill attached. Cloud egress, duplicate storage, and inference on irrelevant tokens all scale with volume, not usefulness — and DataBahn's CEO, Nanda Santhana, put the underlying insight bluntly: "AI is only as good as the enterprise data it can understand."

What does DataBahn actually do?

DataBahn sits between raw enterprise telemetry and whatever consumes it — a SIEM, a data warehouse, or an AI agent — and decides what's worth forwarding. It calls the underlying system AIDI (Autonomous In-Stream Data Intelligence): it analyzes, enriches, and governs data in real time as it moves, instead of after it lands.

The company ingests from more than 600 sources across healthcare, financial services, manufacturing, and transportation clients. Instead of replicating everything into a central lake, it routes a filtered, enriched stream to whichever downstream tool actually needs it — which is where the cost savings come from.

The numbers behind the pitch

MetricFigure
Series B (July 30, 2026)$40M
Total funding to date$59M
Series A (June 2025)$17M
Data sources ingested600+
YoY revenue growth400%+
Net revenue retention180%
Customer churn0%
PoC win rate97%

A 97% proof-of-concept win rate against a field of established observability and data-pipeline vendors is the number that should catch a competitor's attention. It suggests buyers aren't comparing DataBahn on price — they're comparing it on whether the alternative even solves the routing problem at all.

Is this the same problem web scraping solves?

Yes, mirrored. DataBahn cleans and routes data enterprises already own before it reaches a model. ScrapeOps exists because most of what a model needs isn't owned yet — it's scattered across the open web, and the job is turning one query into hundreds of deduplicated, comprehension-ready sources instead of a pile of raw HTML.

Both problems come from the same root cause: AI systems don't fail from too little data. They fail from too much of the wrong data reaching the model at the wrong stage. Insight Partners is betting $40 million that enterprises will keep paying to fix that on the internal side — the external side, the open web, is no less broken.

Publishers issuing cease-and-desist letters to crawlers, Cloudflare defaulting to block AI agents by September 15, and courts still sorting out who owns the right to scrape search results all point the same direction: getting clean data to a model, from any source, now costs real money and real engineering. DataBahn just proved investors will fund that on the enterprise side. The open-web side isn't optional either — it's just less visible on a pitch deck.

Read more from Dekrypt Labs' dispatches or dig into the underlying research.

Frequently Asked Questions

What does DataBahn's $40M Series B fund? The round, led by Insight Partners with Forgepoint, GTM Capital, and S3 Ventures, funds DataBahn's "agentic data control plane" — infrastructure that filters and routes enterprise telemetry to AI systems instead of dumping it all into a central data lake.

How much has DataBahn raised in total? DataBahn has raised $59 million since its 2024 founding: a $17 million Series A in June 2025 led by Forgepoint Capital, followed by the $40 million Series B announced July 30, 2026, led by Insight Partners.

Why do AI data pipelines cost so much? Every duplicated dataset, unfiltered log stream, and irrelevant token forwarded to a model incurs cloud egress, storage, and inference charges. Routing only relevant, enriched data — rather than everything available — cuts those costs without cutting model quality.

How is this different from what ScrapeOps does? DataBahn filters data enterprises already hold internally before it reaches a model. ScrapeOps solves the earlier problem — turning a query into deduplicated, comprehension-ready sources pulled from the open web, where the data isn't owned yet.

Abhishek Gupta is Co-Founder at Dekrypt Labs, building ScrapeOps — the data acquisition engine that turns any question into clean, deduplicated, comprehension-ready sources. dekryptlabs.com