{"jobs":[{"id":"06defc75-e0d5-4326-875f-87a8e2f96c92","title":"Software Engineer, Apps","department":"Engineering","team":"Engineering","employmentType":"FullTime","location":"Palo Alto","secondaryLocations":[],"publishedAt":"2026-07-09T05:35:52.433+00:00","isListed":true,"isRemote":false,"workplaceType":"OnSite","address":{"postalAddress":{"addressRegion":"California","addressCountry":"United States","addressLocality":"Palo Alto"}},"jobUrl":"https://jobs.ashbyhq.com/ollama/06defc75-e0d5-4326-875f-87a8e2f96c92","applyUrl":"https://jobs.ashbyhq.com/ollama/06defc75-e0d5-4326-875f-87a8e2f96c92/application","descriptionHtml":"
Ollama is the most popular way for developers to access open models. What started as an open-source, local-first runtime is now the largest developer network in the open-model ecosystem: 8.9 million monthly active developers and over 67,000+ community-built integrations. We're backed by Y Combinator, Benchmark, 8VC, and Theory Ventures.
Our team is small and talent dense. We're flat, low-ego, and fast-moving. We like people who are truth-seeking, passionate, design-driven, and who enjoy shipping code.
In this role, you’ll work on Ollama’s core user experience: the CLI, Ollama’s app and first-class integrations for Ollama. This involves a strong mix of engineering and design: you'll build new interfaces, ship features end-to-end, and shape how millions of developers experience Ollama.
Build and polish the Ollama desktop app, API, and CLI — the daily driver for millions of developers.
Design and implement new product surfaces for Ollama.
Implement and improve integrations with coding tools, IDEs and more.
You've built a great product that real people use.
You combine strong engineering with an eye for design and product taste.
You can build powerful tools without compromising ease of use.
You're comfortable across the stack — TypeScript/React on the front end, Go or similar on the back, and native apps where it matters.
You make good calls in gray areas, weighing data, UX, and taste.
Bonus: experience with developer tools, AI products, or shipping native/desktop apps.
Ollama is the most popular way for developers to access open models. What started as an open-source, local-first runtime is now the largest developer network in the open-model ecosystem: 8.9 million monthly active developers and over 67,000+ community-built integrations. We're backed by Y Combinator, Benchmark, 8VC, and Theory Ventures.
Our team is small and talent dense. We're flat, low-ego, and fast-moving. We like people who are truth-seeking, passionate, design-driven, and who enjoy shipping code.
You'll work on the heart of Ollama — the local runtime that runs open models on developers' own machines. It loads models, manages memory, drives GPU acceleration across NVIDIA, AMD, Intel, Qualcomm, and Apple Silicon (including our MLX integration), and makes all of it feel instant. You'll work in Go and C/C++ and touch the model formats and inference engines underneath, shipping to macOS, Linux, and Windows across an enormous range of hardware.
Make open models run fast and reliably on consumer and enterprise hardware — from a MacBook Pro to server-grade NVIDIA GPUs.
Own pieces of the runtime: model loading & scheduling memory management, quantization, GPU hardware backends.
Integrate new model architectures and quantization formats so the latest open models work on day one.
Improve cold-start, time-to-first-token, and throughput
Partner with model labs and hardware vendors on early access and deep integrations.
Ship in the open: Ollama is open source, and you'll work with the community
Add support for a new model family end-to-end — format parsing, weights loading, and the defaults that make it useful out of the box.
Cut cold-start for a popular model in half by streaming weights and lazy-loading layers.
Land a new quantization format so a 70B model runs on a single consumer GPU.
Wire up a new GPU backend and find a 2x throughput win with kernel selection and memory tuning.
Improve the \"Auto\" experience — picking the right model and settings for a machine's hardware without the user thinking about it.
You have strong systems fundamentals and are comfortable in Go, C, or C++
You've worked close to the metal — GPU compute, inference, game engines, databases, OS, or networking.
You care about performance and have profiled and optimized real workloads.
You're comfortable shipping to millions of users and handling the long tail of hardware and OS combinations.
Bonus: experience with model quantization, GPU programming (CUDA/Metal/SYCL), Apple MLX
Ollama is the most popular way for developers to access open models. What started as an open-source, local-first runtime is now the largest developer network in the open-model ecosystem: 8.9 million monthly active developers and over 67,000+ community-built integrations. We're backed by Y Combinator, Benchmark, 8VC, and Theory Ventures.
Our team is small and talent dense. We're flat, low-ego, and fast-moving. We like people who are truth-seeking, passionate, design-driven, and who enjoy shipping code.
You'll build Ollama’s cloud, a scalable inference platform that lets developers run large, capable open models in their workflow. You'll work on high-throughput, low-latency distributed systems — inference serving, GPU fleet management, routing, metering, and the platform that Pro, Max, Team, and Enterprise customers rely on to process trillions of tokens.
Build and scale the inference platform that serves every request from ollama.com.
Design the routing and capacity layer that places workloads across GPUs and regions for cost, latency, and availability.
Own multi-tenant infrastructure: isolation, quotas, usage metering, billing, and Pro/Max/team/enterprise tiering.
Build the reliability, observability, and cost controls for our team and customers
You have deep experience with high-throughput, low-latency distributed systems — inference serving, traffic routing, real-time data pipelines, or large-scale APIs.
You're comfortable with cost/performance tradeoffs at scale and have owned a production service end-to-end.
You've worked with Kubernetes, GPU scheduling, or inference infrastructure.
You think in terms of reliability, SLOs, and honest capacity planning.
Bonus: experience building an inference platform, GPU fleet management, or billing/metering for an AI service.
Ollama is the most popular way for developers to access open models. What started as an open-source, local-first runtime is now the largest developer network in the open-model ecosystem: 8.9 million monthly active developers and over 67,000+ community-built integrations. We're backed by Y Combinator, Benchmark, 8VC, and Theory Ventures.
Our team is small and talent dense. We're flat, low-ego, and fast-moving. We like people who are truth-seeking, passionate, design-driven, and who enjoy shipping code.
We're looking for a Product Marketing Manager who can translate deeply technical products into clear, compelling stories that resonate with developers.
This is a highly cross-functional, high-impact role. You'll work closely with product, engineering, and GTM to help define what we build and how we bring it to market. You won't just launch features, you'll influence product direction, define positioning, and help developers understand why Ollama matters.
Own end-to-end go-to-market strategy for new models, features, and major releases
Write launch materials (blog posts, landing pages, demos, videos)
Develop clear, differentiated positioning for Ollama and its features
Continuously refine messaging based on user feedback and market shifts
Help launch day-zero models with research partners
You have 3+ years in product marketing, product management, growth, or a similar role
You have strong technical intuition — you should be an engineer at heart, comfortable understanding developer tools and communicating to a technical audience
You have exceptional communication skills, and the ability to simplify complex concepts
You have high ownership: you can take a project from idea to execution
You find comfort in a fast-moving, ambiguous environment
Ollama is the most popular way for developers to access open models. What started as an open-source, local-first runtime is now the largest developer network in the open-model ecosystem: 8.9 million monthly active developers and over 67,000+ community-built integrations. We're backed by Y Combinator, Benchmark, 8VC, and Theory Ventures.
Our team is small and talent dense. We're flat, low-ego, and fast-moving. We like people who are truth-seeking, passionate, design-driven, and who enjoy shipping code.
You'll be the technical bridge between Ollama and our customers — helping teams deploy Ollama locally and in the cloud, integrate it into their stack, and trust it in regulated environments. You'll work on proofs of concept, reference architectures, onboarding, and ongoing success for the developers building on Ollama. The role balances support engineering — unblocking teams as they deploy and integrate Ollama — with helping Ollama's fastest-growing users succeed.
Build and automate Ollama's support systems
Build reference architectures and deployment guides for specific use cases
Onboard larger teams to Ollama
Be the voice of the customer inside Ollama.
You've been a solutions, customer, pre-sales, or developer-success engineer at a developer-tools, infrastructure, or AI/ML company.
You're hands-on — you can run models, write code, and debug a real deployment.
You're credible with technical buyers and can lead a room of engineers.
You communicate clearly and simplify complex setups for developers
You're comfortable in ambiguity and move fast.
Bonus: experience with model deployment, GPU/ML infrastructure, or regulated-industry compliance.
Ollama is the most popular way for developers to access open models. What started as an open-source, local-first runtime is now the largest developer network in the open-model ecosystem: 8.9 million monthly active developers and over 67,000+ community-built integrations. We're backed by Y Combinator, Benchmark, 8VC, and Theory Ventures.
Our team is small and talent dense. We're flat, low-ego, and fast-moving. We like people who are truth-seeking, passionate, design-driven, and who enjoy shipping code.
You'll be a founding account executive at Ollama. You'll work directly with company’s founders to help growing teams be successful with the product.
Help customers looking to grow with Ollama’s team and enterprise plans be successful
Be the voice of the customer inside Ollama.
Partner with product and engineering on ROI cases, and feed field signal back into pricing, packaging, and features.
Build the sales playbook, collateral, and CRM process from scratch
You've been an account executive at a developer-tools, open-source, infrastructure, or product-led company and closed real enterprise deals.
You're helpful to technical buyers and comfortable in the details of how engineers adopt and deploy software.
You can layer an enterprise motion on bottoms-up adoption while keeping what developers love.
You thrive in ambiguity and move fast — you'd rather ship a v1 motion and iterate.
Ollama is the most popular way for developers to access open models. What started as an open-source, local-first runtime is now the largest developer network in the open-model ecosystem: 8.9 million monthly active developers and over 67,000+ community-built integrations. We're backed by Y Combinator, Benchmark, 8VC, and Theory Ventures.
Our team is small and talent dense. We're flat, low-ego, and fast-moving. We like people who are truth-seeking, passionate, design-driven, and who enjoy shipping code.
You'll build the systems and automation that power Ollama's go-to-market — the connective tissue between product, sales, and customer engineering. This is a hybrid engineering role for someone who thinks like a builder and cares about go-to-market. You'll instrument Ollama’s sales funnel, wire product-usage signals into the sales motion, automate Ollama’s customer funnel, and build the internal tools to help map out customer use cases better.
Instrument the customer funnel from first local install to cloud, team, and enterprise
Wire CRM, product analytics, billing, and communication tools into one coherent GTM stack
Build internal tools and workflows that let GTM move fast without a big team
Run experiments on conversion across the funnel and decide with data
Partner with sales, customer engineering, product, and PMM
You're an engineer who wants to work on revenue, or a GTM-leaning builder who can code
You've built internal tools, data pipelines, or revenue growth systems
You have strong full-stack and data skills
You think in funnels, experiments, and measurement
You're comfortable in ambiguity and ship fast