
Databricks#5
About

Databricks is an American data-and-AI company built around the open-source Apache Spark and the 'lakehouse' architecture, serving enterprises for data engineering, analytics, and AI. It surpassed a $4B+ revenue run-rate in 2026.
From X
Grok now on Databricks
GLM-5.2 is the open-source Claude moment. The demand we’re seeing at Databricks is astonishing. The world is going to see massive adoption of oss LLMs. Also, more companies will shift toward post-training their own models on top of oss models and owning the weights.
Databricks is embracing Chinese open source model GLM-5.2. Earlier, Microsoft said they were looking to embrace DeepSeek. Major American software companies are embracing Chinese open source models to cut token costs. I have long advocated using the Chinese models running in one's own infrastructure. We have Zoho-specific small models that we use in production. This approach also will become common for larger companies, as it becomes easier to set up the training pipeline to train your own org specific model.
Life update: I joined Databricks this week! I thought I’d do another startup after Hyperbolic, but I was surprised by how startup-y Databricks AI is. @alighodsi, @pwendell, @matei_zaharia are in full founder mode. They’re the best founders I’ve met. I like working with people who aren’t “normal” and they definitely aren’t. For example, they invited me to an all-hands before I joined. I’m also impressed by how many former founders are here. @akhilgupta and @hanlintang are incredible leaders. A big bonus: I finally have unlimited Claude Code & Codex tokens! AI adoption on the Databricks AI team is insanely high. Every engineer I’ve met uses AI heavily and shares their own ways to drive agents. Many talented people here. I’m super pumped for this new journey!
Awesome job by the @databricks team My summary: They trained a model called KARL that beats Claude 4.6 and GPT 5.2 on enterprise knowledge tasks (searching docs, cross-referencing info, answering questions over internal data), at ~33% lower cost and ~47% lower latency. The key insight: instead of throwing expensive frontier models at enterprise search, you can use reinforcement learning on synthetic data to train a smaller model that's faster, cheaper, AND better at the specific task. RL went beyond making the model more accurate. I t learned to search more efficiently (fewer wasted queries, better knowing when to stop searching and commit to an answer). They're opening this RL pipeline to Databricks customers so they can build their own custom RL-optimized agents for high-volume workloads. I think we'll continue to see data platforms become agent platforms. Databricks' KARL paper is really an agent platform play. The pitch: you already store your enterprise data in the Lakehouse, now Databricks will train a custom RL agent that searches and reasons over it, tuned specifically for your highest-volume workloads (workloads = apps = agents). The business move is closing the loop: data storage → retrieval → custom agent training → serving, all on Databricks. They're turning "your data lives here" into "your agents live here too." Kudos @alighodsi @matei_zaharia @rxin
Databricks CEO Ali Ghodsi says the company will go public eventually but 2026 is a “terrible year to go public.” https://t.co/J3k2UXmMGi
News
Polls
Will Databricks IPO before Stripe?
Founders








