/00 — boot sequence

Hello.

Article

Ollama Raises $65M Series B: Local AI's Platform Moment Has Arrived

July 10, 2026•5 min read
ollama open-source ai-development local-ai funding developer-tools

Ollama, the open-source tool that lets developers run AI models on their own machines with a single command, has raised a $65 million Series B funding round. Led by Theory Ventures, the round brings Ollama's total funding to $88 million, and the platform now serves 8.9 million monthly active developers. This funding marks a pivotal moment for local AI development and signals that open-weight models are no longer just a research curiosity; they are becoming the default way developers build with AI.

Background: From Docker to AI

Ollama's founding story is deeply tied to the container revolution. Founder Jeff Morgan and co-founder Michael Chiang previously built Kitematic, a tool that made Docker dead-simple to run. When Docker acquired Kitematic in 2015, their work evolved into Docker Desktop, which today serves over ten million developers worldwide.

Ten years later, the same team is applying this playbook to AI. When powerful open models started emerging in 2023, they were geared toward researchers, not programmers. Getting them running locally required wrestling with Python environments, CUDA toolkits, and arcane CLI flags. Morgan and Chiang saw an opportunity to make open models as easy to run as any other piece of software.

Ollama launched in 2023 with a simple value proposition: install the app, run a single command, and you have a local AI model ready to use. No API keys, no cloud credits, no expensive hardware required beyond what most developers already own.

Why This Matters

Ollama's growth is staggering by any measure. The platform now has 176,000 GitHub stars and nearly 17,000 forks, making it one of the fastest-growing open-source developer tools ever. With 67,000+ integrations and adoption by 85% of Fortune 500 companies, Ollama has become the de facto gateway to open-weight models.

The funding comes at a critical inflection point. Around January 2026, when OpenClaw and other agentic models hit the scene, open models suddenly became capable of real agentic tasks like coding. This was the proving point for Ollama as a business, Morgan told TechCrunch. Enterprise interest shifted from "can we use open models?" to "how do we scale open models across our entire engineering organization?"

The Technical Breakdown

Ollama works as a local inference runtime that wraps the complexity of model deployment behind a clean, OpenAI-compatible API. Developers interact with it through a CLI or a GUI, and the same commands work whether you are running a 7B model on your laptop or swapping to a 400B model in Ollama's cloud.

Key technical features include:

  • Local-first design: Models run on your hardware with no data leaving your machine
  • OpenAI-compatible API: Drop-in replacement for existing OpenAI tooling
  • Cross-platform support: macOS, Windows, Linux, with MLX acceleration on Apple Silicon
  • Structured outputs: Constrain model output to JSON schemas
  • Tool calling: Models can use tools and APIs natively
  • Image generation: Experimental support for local image generation

Ollama's cloud inference service extends the local experience. Instead of per-token pricing (which is hard to forecast), Ollama charges based on GPU time, with tiers from free to $100 per month. This model has proven popular: the company reports that cloud token volume has more than doubled every month.

The Competitive Landscape

Ollama operates in a growing ecosystem of local AI tools. Direct competitors include LM Studio, which offers a polished GUI for browsing and chatting with models, and providers like Fireworks, Groq, and Together that offer high-speed cloud inference.

But Ollama's unique advantage is its dual local-and-cloud model. Developers start locally for free, then seamlessly scale to cloud inference when they need larger models. This "same API, same workflow" approach is a direct parallel to what Docker did for containers: abstract away the infrastructure so developers can focus on building.

As Benchmark's Peter Fenton put it: "What Jeff and Michael built with Docker is being used by 10 million-plus developers every day. The creative powers to create a product that goes to ubiquity for developers is extremely rare."

How to Get Started

Getting started with Ollama is straightforward:

bash

For production deployments, Ollama integrates with Docker, Kubernetes, and popular frameworks like LangChain and Continue. You can also use it as a backend for Open WebUI, a self-hosted ChatGPT-like interface with RAG, voice, and plugin support.

The Open vs Closed Model Debate

Ollama's growth reignites the debate about open versus closed AI models. Critics have raised concerns about "enshittification" of developer tools as Ollama builds its cloud business alongside the free open-source project.

But Fenton pushes back on the binary framing: "It's not an either/or. There will be plenty of business for both." Every company with high inference expenses has a vital interest in open models as a cost-control lever. Morgan emphasizes that the cloud service is just an evolution of the same open-source mission: "Nothing has changed for the core product that's free on the desktop."

Frequently Asked Questions

Is Ollama still free? Yes. The core local tool remains completely free and open-source. The cloud service is an optional paid upgrade for developers who need access to larger models.

What models can I run with Ollama? Ollama supports hundreds of models including Llama 3.2, Gemma 4, Qwen3, DeepSeek, Nemotron, GLM, Kimi, MiniMax, and many more. New models are added regularly.

Does Ollama work with VS Code? Yes. Through the Continue extension, you can use Ollama as your coding assistant directly inside VS Code or JetBrains IDEs.

Can I use Ollama in production? Yes. Ollama is used by 85% of Fortune 500 companies in production environments, with Docker, Kubernetes, and cloud deployment options.

How does Ollama make money? Through cloud inference subscriptions (free to $100/month) based on GPU time, not per-token pricing. The local open-source tool remains free.

Key Takeaways

  • Ollama raised $65M Series B led by Theory Ventures, totaling $88M raised
  • Platform now serves 8.9 million monthly developers, growing from zero in under three years
  • 85% of Fortune 500 companies use Ollama for open-weight model inference
  • The company is run by the same team that built Docker Desktop
  • Ollama's dual local-and-cloud model positions it as the platform layer for open-source AI
  • The core local tool remains free and open-source

Conclusion

Ollama's $65M Series B is more than a funding milestone; it is a signal that open-weight AI models have arrived as a mainstream developer tool. What Docker did for containers, Ollama is doing for AI: abstracting away complexity, putting power in developers' hands, and keeping the platform open. For developers who have been waiting for a clear signal that local AI is the future, this is it.


Sources: TechCrunch, SiliconANGLE, Ollama Blog

Automated Transmission

This entry was synthesized and populated dynamically using native API integrations.

Resources & Links