Insight Partners just put $40 million into a company whose entire pitch is that most of the data flowing into your AI models shouldn't be there. DataBahn closed a Series B on July 30, 2026, taking its total raised to $59 million. The bet: AI data pipelines are now expensive enough that filtering what reaches a model is a standalone business.
The short version
Because the cost of feeding an AI model badly has become visible on a P&L. Enterprises spent 2024–2025 dumping every log, event, and telemetry stream into a data lake on the theory that more context makes AI smarter.
That theory has a bill attached. Cloud egress, duplicate storage, and inference on irrelevant tokens all scale with volume, not usefulness — and DataBahn's CEO, Nanda Santhana, put the underlying insight bluntly: "AI is only as good as the enterprise data it can understand."
DataBahn sits between raw enterprise telemetry and whatever consumes it — a SIEM, a data warehouse, or an AI agent — and decides what's worth forwarding. It calls the underlying system AIDI (Autonomous In-Stream Data Intelligence): it analyzes, enriches, and governs data in real time as it moves, instead of after it lands.
The company ingests from more than 600 sources across healthcare, financial services, manufacturing, and transportation clients. Instead of replicating everything into a central lake, it routes a filtered, enriched stream to whichever downstream tool actually needs it — which is where the cost savings come from.
| Metric | Figure |
|---|---|
| Series B (July 30, 2026) | $40M |
| Total funding to date | $59M |
| Series A (June 2025) | $17M |
| Data sources ingested | 600+ |
| YoY revenue growth | 400%+ |
| Net revenue retention | 180% |
| Customer churn | 0% |
| PoC win rate | 97% |
A 97% proof-of-concept win rate against a field of established observability and data-pipeline vendors is the number that should catch a competitor's attention. It suggests buyers aren't comparing DataBahn on price — they're comparing it on whether the alternative even solves the routing problem at all.
Yes, mirrored. DataBahn cleans and routes data enterprises already own before it reaches a model. ScrapeOps exists because most of what a model needs isn't owned yet — it's scattered across the open web, and the job is turning one query into hundreds of deduplicated, comprehension-ready sources instead of a pile of raw HTML.
Both problems come from the same root cause: AI systems don't fail from too little data. They fail from too much of the wrong data reaching the model at the wrong stage. Insight Partners is betting $40 million that enterprises will keep paying to fix that on the internal side — the external side, the open web, is no less broken.
Publishers issuing cease-and-desist letters to crawlers, Cloudflare defaulting to block AI agents by September 15, and courts still sorting out who owns the right to scrape search results all point the same direction: getting clean data to a model, from any source, now costs real money and real engineering. DataBahn just proved investors will fund that on the enterprise side. The open-web side isn't optional either — it's just less visible on a pitch deck.
Read more from Dekrypt Labs' dispatches or dig into the underlying research.
What does DataBahn's $40M Series B fund? The round, led by Insight Partners with Forgepoint, GTM Capital, and S3 Ventures, funds DataBahn's "agentic data control plane" — infrastructure that filters and routes enterprise telemetry to AI systems instead of dumping it all into a central data lake.
How much has DataBahn raised in total? DataBahn has raised $59 million since its 2024 founding: a $17 million Series A in June 2025 led by Forgepoint Capital, followed by the $40 million Series B announced July 30, 2026, led by Insight Partners.
Why do AI data pipelines cost so much? Every duplicated dataset, unfiltered log stream, and irrelevant token forwarded to a model incurs cloud egress, storage, and inference charges. Routing only relevant, enriched data — rather than everything available — cuts those costs without cutting model quality.
How is this different from what ScrapeOps does? DataBahn filters data enterprises already hold internally before it reaches a model. ScrapeOps solves the earlier problem — turning a query into deduplicated, comprehension-ready sources pulled from the open web, where the data isn't owned yet.
Abhishek Gupta is Co-Founder at Dekrypt Labs, building ScrapeOps — the data acquisition engine that turns any question into clean, deduplicated, comprehension-ready sources. dekryptlabs.com