{"jobs":[{"id":"7f934a1b-1845-4f25-8fe4-1def2b039f60","title":"Member of Technical Staff, Performance and Scale","department":"Research & Engineering","team":"Research & Engineering","employmentType":"FullTime","location":"San Francisco","secondaryLocations":[],"publishedAt":"2026-01-22T02:29:35.851+00:00","isListed":true,"isRemote":false,"workplaceType":"OnSite","address":{"postalAddress":{"addressRegion":"California","addressCountry":"United States","addressLocality":"San Francisco"}},"jobUrl":"https://jobs.ashbyhq.com/inferact/7f934a1b-1845-4f25-8fe4-1def2b039f60","applyUrl":"https://jobs.ashbyhq.com/inferact/7f934a1b-1845-4f25-8fe4-1def2b039f60/application","descriptionHtml":"

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware—a position that took years to build.

About the Role

We're looking for an infrastructure engineer to build the distributed systems that power inference at global scale. You'll design and implement the foundational layers that enable vLLM to serve models across thousands of accelerators with minimal latency and maximum reliability. Tomorrow, deploying a frontier model at scale should be as straightforward as spinning up a serverless database. The complexity doesn't disappear as it gets absorbed into the infrastructure you're building.

Skills and Qualifications

Minimum qualifications:

Preferred qualifications:

Bonus points if you have:

Logistics

","descriptionPlain":"Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware—a position that took years to build.\n\n\n\n\nABOUT THE ROLE\n\nWe're looking for an infrastructure engineer to build the distributed systems that power inference at global scale. You'll design and implement the foundational layers that enable vLLM to serve models across thousands of accelerators with minimal latency and maximum reliability. Tomorrow, deploying a frontier model at scale should be as straightforward as spinning up a serverless database. The complexity doesn't disappear as it gets absorbed into the infrastructure you're building.\n\n\n\n\nSKILLS AND QUALIFICATIONS\n\nMinimum qualifications:\n\n - Bachelor's degree or equivalent experience in computer science, engineering, or similar.\n\n - Strong systems programming skills in Rust, Go, or C++.\n\n - Experience designing and building high-performance distributed systems at scale.\n\n - Understanding of network protocols and high-performance I/O.\n\n - Ability to debug complex distributed systems issues.\n\nPreferred qualifications:\n\n - Experience with ML serving infrastructure and disaggregated inference architecture.\n\n - Familiarity with GPU programming models and memory hierarchies.\n\n - Knowledge of GPU interconnects (NVLink, InfiniBand, RoCE) and their performance characteristics.\n\n - Track record of improving system reliability and performance at scale.\n\nBonus points if you have:\n\n - Prior experience in supporting large‑scale model training or inference environments.\n \n \n\n\nLOGISTICS\n\n - Location: This role is based in San Francisco, California. Will consider remote in the US for exceptional candidates.\n\n - Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.\n\n - Visa sponsorship: We sponsor visas on a case-by-case basis.\n\n - Benefits: Inferact offers generous health, dental, and vision benefits as well as 401(k) company match."},{"id":"384d9db8-c712-4caa-8091-444b4189e161","title":"Member of Technical Staff, Kernel Engineering","department":"Research & Engineering","team":"Research & Engineering","employmentType":"FullTime","location":"San Francisco","secondaryLocations":[],"publishedAt":"2026-01-22T02:31:17.947+00:00","isListed":true,"isRemote":false,"workplaceType":"OnSite","address":{"postalAddress":{"addressRegion":"California","addressCountry":"United States","addressLocality":"San Francisco"}},"jobUrl":"https://jobs.ashbyhq.com/inferact/384d9db8-c712-4caa-8091-444b4189e161","applyUrl":"https://jobs.ashbyhq.com/inferact/384d9db8-c712-4caa-8091-444b4189e161/application","descriptionHtml":"

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware—a position that took years to build.

About the Role

We're looking for a performance engineer to squeeze every FLOP out of modern accelerators. You'll write the kernels and low-level optimizations that make vLLM the fastest inference engine in the world. Your code will run on hundreds of accelerator types, from NVIDIA GPUs to emerging silicon. When hardware vendors develop new chips, they integrate with vLLM. You'll work directly with these teams to ensure we're extracting maximum performance from every generation of hardware.

Skills and Qualifications

Minimum qualifications:

Preferred qualifications:

Bonus points if you have:

Logistics

","descriptionPlain":"Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware—a position that took years to build.\n\n\n\n\nABOUT THE ROLE\n\nWe're looking for a performance engineer to squeeze every FLOP out of modern accelerators. You'll write the kernels and low-level optimizations that make vLLM the fastest inference engine in the world. Your code will run on hundreds of accelerator types, from NVIDIA GPUs to emerging silicon. When hardware vendors develop new chips, they integrate with vLLM. You'll work directly with these teams to ensure we're extracting maximum performance from every generation of hardware.\n\n\n\n\nSKILLS AND QUALIFICATIONS\n\nMinimum qualifications:\n\n - Bachelor's degree or equivalent experience in computer science, engineering, or similar.\n\n - Deep experience writing CUDA kernels or equivalent (CuTeDSL, Triton, TileLang, Pallas).\n\n - Strong understanding of GPU architecture: memory hierarchy, warp scheduling, tiling, tensor cores.\n\n - Proficiency in C++ and Python with demonstrated ability to write high-performance code.\n\n - Experience with profiling tools (Nsight, rocprof) and performance optimization methodologies.\n\n - Obsession with benchmarks and squeezing every percentage point of speedup.\n\nPreferred qualifications:\n\n - Experience with ML-specific kernel optimization (FlashAttention, fused kernels).\n\n - Knowledge of quantization techniques (INT8, FP8, mixed-precision).\n\n - Familiarity with multiple accelerator platforms (NVIDIA, AMD, TPU, Intel).\n\n - Experience with compiler technologies (LLVM, MLIR, XLA).\n\nBonus points if you have:\n\n - Kernel-related contributions to vLLM or other inference engine projects.\n\n - Contributions to open-source GPU, ML systems, or compiler optimization projects\n\n - Written deep technical blogs on GPU optimization.\n \n \n\n\nLOGISTICS\n\n - Location: This role is based in San Francisco, California. Will consider remote in the US for exceptional candidates.\n\n - Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.\n\n - Visa sponsorship: We sponsor visas on a case-by-case basis.\n\n - Benefits: Inferact offers generous health, dental, and vision benefits as well as 401(k) company match."},{"id":"3fc0dd9b-14d5-4068-b597-7ec2014e07c1","title":"Member of Technical Staff, Cloud Orchestration","department":"Research & Engineering","team":"Research & Engineering","employmentType":"FullTime","location":"San Francisco","secondaryLocations":[],"publishedAt":"2026-01-22T02:33:27.212+00:00","isListed":true,"isRemote":false,"workplaceType":"OnSite","address":{"postalAddress":{"addressRegion":"California","addressCountry":"United States","addressLocality":"San Francisco"}},"jobUrl":"https://jobs.ashbyhq.com/inferact/3fc0dd9b-14d5-4068-b597-7ec2014e07c1","applyUrl":"https://jobs.ashbyhq.com/inferact/3fc0dd9b-14d5-4068-b597-7ec2014e07c1/application","descriptionHtml":"

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware—a position that took years to build.

About the Role

We're looking for an cloud orchestration engineer to build the operational backbone that keeps vLLM running reliably at massive scale. You'll design the systems for cluster management, deployment automation, and production monitoring that enable teams worldwide to serve AI models without friction. You'll ensure that vLLM deployments are observable, debuggable, and recoverable, turning operational complexity into infrastructure that just works.

Skills and Qualifications

Minimum qualifications:

Preferred qualifications:

Bonus points if you have:

Logistics

","descriptionPlain":"Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware—a position that took years to build.\n\n\n\n\nABOUT THE ROLE\n\nWe're looking for an cloud orchestration engineer to build the operational backbone that keeps vLLM running reliably at massive scale. You'll design the systems for cluster management, deployment automation, and production monitoring that enable teams worldwide to serve AI models without friction. You'll ensure that vLLM deployments are observable, debuggable, and recoverable, turning operational complexity into infrastructure that just works.\n\n\n\n\nSKILLS AND QUALIFICATIONS\n\nMinimum qualifications:\n\n - Bachelor's degree or equivalent experience in computer science, engineering, or similar.\n\n - Strong experience with Kubernetes and container orchestration at scale.\n\n - Experience designing and implementing custom Kubernetes operators.\n\n - Proficiency in Python/Rust/Go and infrastructure-as-code tools (Terraform, Helm, etc).\n\n - Experience managing GPU clusters and debugging hardware issues.\n\n - Ability to work across cloud platforms (AWS, GCP, Azure) and on-premise infrastructure.\n\nPreferred qualifications:\n\n - Experience with ML-specific orchestration tools (Ray, Slurm).\n\n - Knowledge of GPU scheduling, multi-tenancy, and resource optimization.\n\n - Familiarity with vLLM deployment patterns and configuration.\n\n - Track record of improving operational reliability for ML systems.\n\nBonus points if you have:\n\n - Experience deploying inference systems on large-scale GPU (1,000+) clusters.\n \n \n\n\nLOGISTICS\n\n - Location: This role is based in San Francisco, California. Will consider remote in the US for exceptional candidates.\n\n - Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.\n\n - Visa sponsorship: We sponsor visas on a case-by-case basis.\n\n - Benefits: Inferact offers generous health, dental, and vision benefits as well as 401(k) company match."},{"id":"f0d2619d-28e0-4b25-8d30-3ac555071abb","title":"Member of Technical Staff, Exceptional Generalist (Remote)","department":"Research & Engineering","team":"Research & Engineering","employmentType":"FullTime","location":"Remote","secondaryLocations":[],"publishedAt":"2026-01-22T09:56:22.641+00:00","isListed":true,"isRemote":true,"workplaceType":"Remote","address":{"postalAddress":{"addressCountry":"US"}},"jobUrl":"https://jobs.ashbyhq.com/inferact/f0d2619d-28e0-4b25-8d30-3ac555071abb","applyUrl":"https://jobs.ashbyhq.com/inferact/f0d2619d-28e0-4b25-8d30-3ac555071abb/application","descriptionHtml":"

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware—a position that took years to build.

About the Role

This is a globally remote opportunity. We're seeking exceptional generalist engineers who can work across the entire vLLM stack: from low-level GPU kernels to high-level distributed systems. This role is designed for self-directed, autonomous individuals who can identify the highest-leverage problems and solve them end-to-end without constant guidance.

You'll work asynchronously with our San Francisco headquarters while maintaining full ownership of critical infrastructure. You might be optimizing CUDA kernels one week, designing distributed orchestration systems the next, and implementing new model architectures the week after. The work you do will directly impact how the world runs AI inference.

Potential focus areas include:

What We're Looking For

We're looking for engineers who thrive with autonomy. You should be able to take a vague problem statement and turn it into shipped code with minimal supervision. You communicate proactively, over-communicate context across time zones, and know when to ask for help versus when to push forward independently.

Core Requirements:

Technical Depth (strong in at least two):

Preferred Qualifications:

Logistics

","descriptionPlain":"Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware—a position that took years to build.\n\n\n\n\nABOUT THE ROLE\n\nThis is a globally remote opportunity. We're seeking exceptional generalist engineers who can work across the entire vLLM stack: from low-level GPU kernels to high-level distributed systems. This role is designed for self-directed, autonomous individuals who can identify the highest-leverage problems and solve them end-to-end without constant guidance.\n\nYou'll work asynchronously with our San Francisco headquarters while maintaining full ownership of critical infrastructure. You might be optimizing CUDA kernels one week, designing distributed orchestration systems the next, and implementing new model architectures the week after. The work you do will directly impact how the world runs AI inference.\n\nPotential focus areas include:\n\n - Inference Runtime: Push the boundaries of LLM and diffusion model serving. Work at the core of vLLM to optimize how models execute across diverse hardware and architectures.\n\n - Kernel Engineering: Write the low-level kernels and optimizations that make vLLM the fastest inference engine in the world, running on hundreds of accelerator types.\n\n - Performance & Scale: Build the distributed systems that power inference at global scale—design foundational layers enabling vLLM to serve models across thousands of accelerators with minimal latency.\n\n - Cloud Orchestration: Build the operational backbone for cluster management, deployment automation, and production monitoring that enables teams worldwide to serve AI models without friction.\n \n \n\n\nWHAT WE'RE LOOKING FOR\n\nWe're looking for engineers who thrive with autonomy. You should be able to take a vague problem statement and turn it into shipped code with minimal supervision. You communicate proactively, over-communicate context across time zones, and know when to ask for help versus when to push forward independently.\n\n\n\nCore Requirements:\n\n - Bachelor's degree or equivalent experience in computer science, engineering, or similar\n\n - Demonstrated ability to work autonomously and drive projects to completion without close supervision\n\n - Excellent asynchronous communication skills and ability to collaborate effectively across time zones\n\n - Strong track record of shipping high-impact work in complex technical environments\n\n - Deep expertise in at least one of: systems programming, GPU/accelerator programming, distributed systems, or ML infrastructure\n\nTechnical Depth (strong in at least two):\n\n - CUDA kernels or equivalent (Triton, TileLang, Pallas) with deep understanding of GPU architecture\n\n - High-performance distributed systems in Rust, Go, or C++\n\n - Python with PyTorch internals and LLM inference systems (vLLM, TensorRT-LLM, SGLang)\n\n - Kubernetes, container orchestration, and infrastructure-as-code at scale\n\n - Transformer architectures, KV-cache memory management, and model serving\n\nPreferred Qualifications:\n\n - Contributions to vLLM or other major open-source ML/systems projects\n\n - Experience with multiple accelerator platforms (NVIDIA, AMD, TPU, Intel)\n\n - Knowledge of quantization techniques, ML-specific kernel optimization, or compiler technologies\n\n - Track record of improving system reliability and performance at scale\n\n - Written widely-shared technical blogs or impactful side projects in the ML infrastructure space\n \n \n\n\nLOGISTICS\n\n - Location: Fully remote, worldwide. We're timezone-flexible but expect regular overlap with Pacific Time for critical syncs.\n\n - Compensation: We offer competitive compensations (salary + equity) compared to the local market conditions.\n\n - Visa sponsorship: We sponsor visas on a case-by-case basis.\n\n - Benefits: Inferact offers competitive benefits appropriate to your location, including health coverage where applicable.\n\n\n\n"},{"id":"24ea1266-bc29-4838-9a61-8adc1d5bb2c6","title":"Member of Technical Staff, TPU Performance Engineering","department":"Research & Engineering","team":"Research & Engineering","employmentType":"FullTime","location":"San Francisco","secondaryLocations":[],"publishedAt":"2026-06-17T22:25:49.084+00:00","isListed":true,"isRemote":false,"workplaceType":"OnSite","address":{"postalAddress":{"addressRegion":"California","addressCountry":"United States","addressLocality":"San Francisco"}},"jobUrl":"https://jobs.ashbyhq.com/inferact/24ea1266-bc29-4838-9a61-8adc1d5bb2c6","applyUrl":"https://jobs.ashbyhq.com/inferact/24ea1266-bc29-4838-9a61-8adc1d5bb2c6/application","descriptionHtml":"

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.

About the Role

We're looking for a TPU performance engineer to make vLLM a first-class inference engine on Google TPUs. You'll build and optimize TPU backends, compiler integrations, runtime paths, and benchmarking infrastructure using JAX, XLA, Pallas, and related tooling so vLLM can deliver frontier inference performance on TPU hardware.

You'll work at the boundary of inference systems, kernels, compilers, and hardware architecture, improving production-relevant model serving on TPU with clear correctness, latency, and throughput benchmarks. Your work will help make TPU support in vLLM usable, fast, benchmarked, and maintainable.

 

Skills and Qualifications

Minimum qualifications:

Preferred qualifications:

Bonus points if you have:

Logistics

","descriptionPlain":"Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.\n\n\n\nAbout the Role\n\nWe're looking for a TPU performance engineer to make vLLM a first-class inference engine on Google TPUs. You'll build and optimize TPU backends, compiler integrations, runtime paths, and benchmarking infrastructure using JAX, XLA, Pallas, and related tooling so vLLM can deliver frontier inference performance on TPU hardware.\n\nYou'll work at the boundary of inference systems, kernels, compilers, and hardware architecture, improving production-relevant model serving on TPU with clear correctness, latency, and throughput benchmarks. Your work will help make TPU support in vLLM usable, fast, benchmarked, and maintainable.\n\n \n\nSkills and Qualifications\n\nMinimum qualifications:\n\n - Bachelor's degree or equivalent experience in computer science, engineering, systems, machine learning, or similar.\n\n - Hands-on experience building or optimizing TPU workloads using JAX, XLA, Pallas, or related compiler and runtime tooling.\n\n - Deep understanding of TPU execution, memory behavior, compilation, and performance constraints for ML workloads.\n\n - Experience optimizing ML kernels or inference paths such as attention, GEMM, sampling, KV cache, fused kernels, or backend runtime paths.\n\n - Strong performance profiling and benchmarking skills, with the ability to use measurements, compiler artifacts, correctness tests, and reproducible benchmarks to guide optimization work.\n\nPreferred qualifications:\n\n - Experience with vLLM, SGLang, TensorRT-LLM, XLA-based serving, or other LLM inference systems.\n\n - Familiarity with batching, KV cache, decoding, serving tradeoffs, and backend performance constraints in production inference systems.\n\n - Experience with compiler technologies such as XLA, MLIR, LLVM, Pallas, or other kernel DSLs, including lowering, fusion, and backend code generation.\n\n - Knowledge of quantization methods such as INT8, FP8, mixed precision, or TPU-specific numeric formats, including accuracy and performance tradeoffs.\n\nBonus points if you have:\n\n - Contributed to vLLM, JAX/XLA, Pallas, PyTorch/XLA, compiler projects, or other open-source ML infrastructure.\n\n - Built TPU benchmarking infrastructure or automated performance regression detection for accelerator workloads.\n\n - Worked directly with Google TPU ecosystem stakeholders, accelerator platform teams, or early-access programs to ship backend, compiler, or inference performance improvements.\n\nLogistics\n\n - Location: This role is based in San Francisco, California. Will consider remote in the US for exceptional candidates.\n\n - Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.\n\n - Visa sponsorship: We sponsor visas on a case-by-case basis.\n\n - Benefits: Inferact offers generous health, dental, and vision benefits as well as 401(k) company match."},{"id":"e1a91db5-1cd4-4688-863b-33ab88b40a4d","title":"Member of Technical Staff, Developer Relations","department":"Research & Engineering","team":"Research & Engineering","employmentType":"FullTime","location":"San Francisco","secondaryLocations":[],"publishedAt":"2026-06-17T22:39:32.315+00:00","isListed":true,"isRemote":false,"workplaceType":"OnSite","address":{"postalAddress":{"addressRegion":"California","addressCountry":"United States","addressLocality":"San Francisco"}},"jobUrl":"https://jobs.ashbyhq.com/inferact/e1a91db5-1cd4-4688-863b-33ab88b40a4d","applyUrl":"https://jobs.ashbyhq.com/inferact/e1a91db5-1cd4-4688-863b-33ab88b40a4d/application","descriptionHtml":"

Overview

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.

About the Role

We're looking for a Developer Relations Engineer to help make vLLM the default way developers understand, build, and scale AI inference. This is not a generic DevRel role. We're looking for a inference systems educator-builder: someone who can understand vLLM as a deep LLM inference systems project, teach hard technical concepts clearly, and create public artifacts that help practitioners build better systems.

You'll write technical deep dives, build demos, create tutorials, contribute to docs and examples, host workshops, and help developers understand topics like KV cache, continuous batching, prefix caching, prefill and decode, quantization, GPU serving, latency versus throughput, and model-server tradeoffs across vLLM and adjacent systems. Your work will shape how the broader AI infrastructure community learns, adopts, and builds with vLLM.

Skills and Qualifications

Minimum qualifications:

Preferred qualifications:

Bonus points if you have:

Logistics

","descriptionPlain":"Overview\n\nInferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.\n\nAbout the Role\n\nWe're looking for a Developer Relations Engineer to help make vLLM the default way developers understand, build, and scale AI inference. This is not a generic DevRel role. We're looking for a inference systems educator-builder: someone who can understand vLLM as a deep LLM inference systems project, teach hard technical concepts clearly, and create public artifacts that help practitioners build better systems.\n\nYou'll write technical deep dives, build demos, create tutorials, contribute to docs and examples, host workshops, and help developers understand topics like KV cache, continuous batching, prefix caching, prefill and decode, quantization, GPU serving, latency versus throughput, and model-server tradeoffs across vLLM and adjacent systems. Your work will shape how the broader AI infrastructure community learns, adopts, and builds with vLLM.\n\n\n\nSkills and Qualifications\n\nMinimum qualifications:\n\n - Bachelor's degree or equivalent experience in computer science, engineering, machine learning, systems, or similar.\n\n - Strong technical understanding of LLM inference systems, model serving, GPU inference, distributed runtimes, scheduling, batching, quantization, or related infrastructure.\n\n - Ability to credibly explain systems concepts such as KV cache, PagedAttention, continuous batching, prefill / decode scheduling, prefix caching, speculative decoding, tensor parallelism, data parallelism, or latency versus throughput tradeoffs.\n\n - Experience with vLLM or adjacent inference technologies such as SGLang, TensorRT-LLM, TGI, LoRAX, Ray Serve, FlashInfer, BentoML, Baseten-style serving platforms, or similar systems.\n\n - A strong public portfolio of technical artifacts, such as blogs, tutorials, workshops, courses, OSS docs, benchmark posts, architecture explainers, conference talks, demos, or runnable repositories.\n\n - Ability to write and teach for practitioners without sounding like a content marketer.\n\n - Strong engineering judgment, product taste, and ability to turn raw technical material into useful developer education.\n\nPreferred qualifications:\n\n - Prior work in ML systems, distributed systems, HPC, compilers, GPU kernels, serving infrastructure, MLOps, developer tooling, or open-source infrastructure.\n\n - Experience creating technical content that teaches reusable mental models, not just product features.\n\n - Experience contributing to developer-facing open source through docs, tutorials, examples, cookbooks, demos, or community support.\n\n - Existing credibility or community presence in AI infrastructure, OSS, CUDA / GPU, Ray, vLLM, PyTorch, Modal, BentoML, Baseten, Predibase, Together AI, Anyscale, LMSYS, or similar ecosystems.\n\n - Ability to host workshops, create hands-on labs, present technical talks, and help developers move from concept to working code.\n\nBonus points if you have:\n\n - Written widely-shared technical blogs, courses, or architecture deep dives on LLM inference, model serving, GPU serving, or ML systems.\n\n - Built demos, benchmarks, tutorials, or repositories around vLLM, SGLang, TensorRT-LLM, TGI, Ray Serve, FlashInfer, or related systems.\n\n - Contributed to open-source ML infrastructure, inference systems, developer tooling, or technical education projects.\n\n - Created practitioner-facing content with code, diagrams, benchmarks, demos, or end-to-end labs.\n\n - Built a durable personal portfolio that demonstrates technical depth, taste, and a strong point of view.\n\n\n\nLogistics\n\n - Location: This role is based in San Francisco, California. Will consider remote in the US for exceptional candidates.\n\n - Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.\n\n - Visa sponsorship: We sponsor visas on a case-by-case basis.\n\n - Benefits: Inferact offers generous health, dental, and vision benefits as well as 401(k) company match."},{"id":"43c0ca54-fcf5-41fa-83a1-38800c75ccc0","title":"Member of Technical Staff, Inference ","department":"Research & Engineering","team":"Research & Engineering","employmentType":"FullTime","location":"San Francisco","secondaryLocations":[],"publishedAt":"2026-06-18T18:55:18.703+00:00","isListed":true,"isRemote":false,"workplaceType":"OnSite","address":{"postalAddress":{"addressRegion":"California","addressCountry":"United States","addressLocality":"San Francisco"}},"jobUrl":"https://jobs.ashbyhq.com/inferact/43c0ca54-fcf5-41fa-83a1-38800c75ccc0","applyUrl":"https://jobs.ashbyhq.com/inferact/43c0ca54-fcf5-41fa-83a1-38800c75ccc0/application","descriptionHtml":"

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware—a position that took years to build.

About the Role

We're looking for an inference runtime engineer to push the boundaries of what's possible in LLM and diffusion model serving. Models grow larger. Architectures shift: mixture-of-experts, multimodal, agentic. Every breakthrough demands innovations on the inference engine itself. You'll work at the core of vLLM, optimizing how models execute across diverse hardware and architectures. Your work will directly impact how the world runs AI inference.

Skills and Qualifications

Minimum qualifications:

Preferred qualifications:

Bonus points if you have:

Logistics

","descriptionPlain":"Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware—a position that took years to build.\n\n\n\n\nABOUT THE ROLE\n\nWe're looking for an inference runtime engineer to push the boundaries of what's possible in LLM and diffusion model serving. Models grow larger. Architectures shift: mixture-of-experts, multimodal, agentic. Every breakthrough demands innovations on the inference engine itself. You'll work at the core of vLLM, optimizing how models execute across diverse hardware and architectures. Your work will directly impact how the world runs AI inference. \n\n\n\n\nSKILLS AND QUALIFICATIONS\n\nMinimum qualifications:\n\n - Bachelor's degree or equivalent experience in computer science, engineering, or similar.\n\n - Deep understanding of transformer architectures and their variants.\n\n - Strong programming skills in Python with experience in PyTorch internals.\n\n - Experience with LLM inference systems (vLLM, TensorRT-LLM, SGLang, TGI).\n\n - Ability to read and implement model architectures and inference techniques from research papers.\n\n - Demonstrate the ability to contribute performant and maintainable code and debug in complex ML codebases.\n\nPreferred qualifications:\n\n - Deep understanding of KV-cache memory management, prefix caching, and hybrid model serving.\n\n - Familiarity with RL frameworks and algorithms for LLMs.\n\n - Experience with multimodal inference (audio/image/video/text).\n\n - Contributions to open-source ML or system infrastructure projects.\n\nBonus points if you have:\n\n - Implemented core features in vLLM or other inference engine projects.\n\n - Contributed to vLLM integrations (verl, OpenRLHF, Unsloth, LlamaFactory, etc).\n\n - Written widely-shared technical blogs or side projects on vLLM or LLM inference.\n \n \n\n\nLOGISTICS\n\n - Location: This role is based in San Francisco, California. Will consider remote in the US for exceptional candidates.\n\n - Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.\n\n - Visa sponsorship: We sponsor visas on a case-by-case basis.\n\n - Benefits: Inferact offers generous health, dental, and vision benefits as well as 401(k) company match."},{"id":"1f7ff2b3-bcbf-46e3-abcc-ada332250a64","title":"Member of Technical Staff, Cloud Orchestration","department":"Research & Engineering","team":"Research & Engineering","employmentType":"FullTime","location":"Singapore","secondaryLocations":[],"publishedAt":"2026-06-24T05:24:22.610+00:00","isListed":true,"isRemote":false,"workplaceType":"OnSite","address":{"postalAddress":{"addressRegion":"Singapore","addressCountry":"Singapore","addressLocality":"Singapore"}},"jobUrl":"https://jobs.ashbyhq.com/inferact/1f7ff2b3-bcbf-46e3-abcc-ada332250a64","applyUrl":"https://jobs.ashbyhq.com/inferact/1f7ff2b3-bcbf-46e3-abcc-ada332250a64/application","descriptionHtml":"

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware—a position that took years to build.

About the Role

We're looking for an cloud orchestration engineer to build the operational backbone that keeps vLLM running reliably at massive scale. You'll design the systems for cluster management, deployment automation, and production monitoring that enable teams worldwide to serve AI models without friction. You'll ensure that vLLM deployments are observable, debuggable, and recoverable, turning operational complexity into infrastructure that just works.

Skills and Qualifications

Minimum qualifications:

Preferred qualifications:

Bonus points if you have:

Logistics

","descriptionPlain":"Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware—a position that took years to build.\n\n\n\n\nABOUT THE ROLE\n\nWe're looking for an cloud orchestration engineer to build the operational backbone that keeps vLLM running reliably at massive scale. You'll design the systems for cluster management, deployment automation, and production monitoring that enable teams worldwide to serve AI models without friction. You'll ensure that vLLM deployments are observable, debuggable, and recoverable, turning operational complexity into infrastructure that just works.\n\n\n\n\nSKILLS AND QUALIFICATIONS\n\nMinimum qualifications:\n\n - Bachelor's degree or equivalent experience in computer science, engineering, or similar.\n\n - Strong experience with Kubernetes and container orchestration at scale.\n\n - Experience designing and implementing custom Kubernetes operators.\n\n - Proficiency in Python/Rust/Go and infrastructure-as-code tools (Terraform, Helm, etc).\n\n - Experience managing GPU clusters and debugging hardware issues.\n\n - Ability to work across cloud platforms (AWS, GCP, Azure) and on-premise infrastructure.\n\nPreferred qualifications:\n\n - Experience with ML-specific orchestration tools (Ray, Slurm).\n\n - Knowledge of GPU scheduling, multi-tenancy, and resource optimization.\n\n - Familiarity with vLLM deployment patterns and configuration.\n\n - Track record of improving operational reliability for ML systems.\n\nBonus points if you have:\n\n - Experience deploying inference systems on large-scale GPU (1,000+) clusters.\n \n \n\n\nLOGISTICS\n\n - Location: This role is based in Singapore.\n\n - Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is S$200,000 to S$400,000 annually + equity.\n\n - Visa sponsorship: We sponsor visas on a case-by-case basis.\n\n - Benefits: Inferact offers a generous benefits package, including medical, dental, and vision coverage."},{"id":"bc3e42e7-0a71-438a-9815-67a88d5a1efc","title":"Member of Technical Staff, Inference","department":"Research & Engineering","team":"Research & Engineering","employmentType":"FullTime","location":"Singapore","secondaryLocations":[],"publishedAt":"2026-06-24T05:27:33.051+00:00","isListed":true,"isRemote":false,"workplaceType":"OnSite","address":{"postalAddress":{"addressRegion":"Singapore","addressCountry":"Singapore","addressLocality":"Singapore"}},"jobUrl":"https://jobs.ashbyhq.com/inferact/bc3e42e7-0a71-438a-9815-67a88d5a1efc","applyUrl":"https://jobs.ashbyhq.com/inferact/bc3e42e7-0a71-438a-9815-67a88d5a1efc/application","descriptionHtml":"

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware—a position that took years to build.

About the Role

We're looking for an inference runtime engineer to push the boundaries of what's possible in LLM and diffusion model serving. Models grow larger. Architectures shift: mixture-of-experts, multimodal, agentic. Every breakthrough demands innovations on the inference engine itself. You'll work at the core of vLLM, optimizing how models execute across diverse hardware and architectures. Your work will directly impact how the world runs AI inference.

Skills and Qualifications

Minimum qualifications:

Preferred qualifications:

Bonus points if you have:

Logistics

","descriptionPlain":"Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware—a position that took years to build.\n\n\n\n\nABOUT THE ROLE\n\nWe're looking for an inference runtime engineer to push the boundaries of what's possible in LLM and diffusion model serving. Models grow larger. Architectures shift: mixture-of-experts, multimodal, agentic. Every breakthrough demands innovations on the inference engine itself. You'll work at the core of vLLM, optimizing how models execute across diverse hardware and architectures. Your work will directly impact how the world runs AI inference. \n\n\n\n\nSKILLS AND QUALIFICATIONS\n\nMinimum qualifications:\n\n - Bachelor's degree or equivalent experience in computer science, engineering, or similar.\n\n - Deep understanding of transformer architectures and their variants.\n\n - Strong programming skills in Python with experience in PyTorch internals.\n\n - Experience with LLM inference systems (vLLM, TensorRT-LLM, SGLang, TGI).\n\n - Ability to read and implement model architectures and inference techniques from research papers.\n\n - Demonstrate the ability to contribute performant and maintainable code and debug in complex ML codebases.\n\nPreferred qualifications:\n\n - Deep understanding of KV-cache memory management, prefix caching, and hybrid model serving.\n\n - Familiarity with RL frameworks and algorithms for LLMs.\n\n - Experience with multimodal inference (audio/image/video/text).\n\n - Contributions to open-source ML or system infrastructure projects.\n\nBonus points if you have:\n\n - Implemented core features in vLLM or other inference engine projects.\n\n - Contributed to vLLM integrations (verl, OpenRLHF, Unsloth, LlamaFactory, etc).\n\n - Written widely-shared technical blogs or side projects on vLLM or LLM inference.\n \n \n\n\nLOGISTICS\n\n - Location: This role is based in Singapore.\n\n - Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is S$200,000 to S$400,000 annually + equity.\n\n - Visa sponsorship: We sponsor visas on a case-by-case basis.\n\n - Benefits: Inferact offers a generous benefits package, including medical, dental, and vision coverage."},{"id":"9c5be5ae-2269-4061-82a1-4c9903c25355","title":"Member of Technical Staff, Kernel Engineering","department":"Research & Engineering","team":"Research & Engineering","employmentType":"FullTime","location":"Singapore","secondaryLocations":[],"publishedAt":"2026-06-24T05:29:46.412+00:00","isListed":true,"isRemote":false,"workplaceType":"OnSite","address":{"postalAddress":{"addressRegion":"Singapore","addressCountry":"Singapore","addressLocality":"Singapore"}},"jobUrl":"https://jobs.ashbyhq.com/inferact/9c5be5ae-2269-4061-82a1-4c9903c25355","applyUrl":"https://jobs.ashbyhq.com/inferact/9c5be5ae-2269-4061-82a1-4c9903c25355/application","descriptionHtml":"

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware—a position that took years to build.

About the Role

We're looking for a performance engineer to squeeze every FLOP out of modern accelerators. You'll write the kernels and low-level optimizations that make vLLM the fastest inference engine in the world. Your code will run on hundreds of accelerator types, from NVIDIA GPUs to emerging silicon. When hardware vendors develop new chips, they integrate with vLLM. You'll work directly with these teams to ensure we're extracting maximum performance from every generation of hardware.

Skills and Qualifications

Minimum qualifications:

Preferred qualifications:

Bonus points if you have:

Logistics

","descriptionPlain":"Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware—a position that took years to build.\n\n\n\n\nABOUT THE ROLE\n\nWe're looking for a performance engineer to squeeze every FLOP out of modern accelerators. You'll write the kernels and low-level optimizations that make vLLM the fastest inference engine in the world. Your code will run on hundreds of accelerator types, from NVIDIA GPUs to emerging silicon. When hardware vendors develop new chips, they integrate with vLLM. You'll work directly with these teams to ensure we're extracting maximum performance from every generation of hardware.\n\n\n\n\nSKILLS AND QUALIFICATIONS\n\nMinimum qualifications:\n\n - Bachelor's degree or equivalent experience in computer science, engineering, or similar.\n\n - Deep experience writing CUDA kernels or equivalent (CuTeDSL, Triton, TileLang, Pallas).\n\n - Strong understanding of GPU architecture: memory hierarchy, warp scheduling, tiling, tensor cores.\n\n - Proficiency in C++ and Python with demonstrated ability to write high-performance code.\n\n - Experience with profiling tools (Nsight, rocprof) and performance optimization methodologies.\n\n - Obsession with benchmarks and squeezing every percentage point of speedup.\n\nPreferred qualifications:\n\n - Experience with ML-specific kernel optimization (FlashAttention, fused kernels).\n\n - Knowledge of quantization techniques (INT8, FP8, mixed-precision).\n\n - Familiarity with multiple accelerator platforms (NVIDIA, AMD, TPU, Intel).\n\n - Experience with compiler technologies (LLVM, MLIR, XLA).\n\nBonus points if you have:\n\n - Kernel-related contributions to vLLM or other inference engine projects.\n\n - Contributions to open-source GPU, ML systems, or compiler optimization projects\n\n - Written deep technical blogs on GPU optimization.\n \n \n\n\nLOGISTICS\n\n - Location: This role is based in Singapore.\n\n - Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is S$200,000 to S$400,000 annually + equity.\n\n - Visa sponsorship: We sponsor visas on a case-by-case basis.\n\n - Benefits: Inferact offers a generous benefits package, including medical, dental, and vision coverage."},{"id":"c5c5b219-4949-44ed-8768-6a7bc09f28dd","title":"Member of Technical Staff, Performance and Scale","department":"Research & Engineering","team":"Research & Engineering","employmentType":"FullTime","location":"Singapore","secondaryLocations":[],"publishedAt":"2026-06-24T05:30:20.019+00:00","isListed":true,"isRemote":false,"workplaceType":"OnSite","address":{"postalAddress":{"addressRegion":"Singapore","addressCountry":"Singapore","addressLocality":"Singapore"}},"jobUrl":"https://jobs.ashbyhq.com/inferact/c5c5b219-4949-44ed-8768-6a7bc09f28dd","applyUrl":"https://jobs.ashbyhq.com/inferact/c5c5b219-4949-44ed-8768-6a7bc09f28dd/application","descriptionHtml":"

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware—a position that took years to build.

About the Role

We're looking for an infrastructure engineer to build the distributed systems that power inference at global scale. You'll design and implement the foundational layers that enable vLLM to serve models across thousands of accelerators with minimal latency and maximum reliability. Tomorrow, deploying a frontier model at scale should be as straightforward as spinning up a serverless database. The complexity doesn't disappear as it gets absorbed into the infrastructure you're building.

Skills and Qualifications

Minimum qualifications:

Preferred qualifications:

Bonus points if you have:

Logistics

","descriptionPlain":"Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware—a position that took years to build.\n\n\n\n\nABOUT THE ROLE\n\nWe're looking for an infrastructure engineer to build the distributed systems that power inference at global scale. You'll design and implement the foundational layers that enable vLLM to serve models across thousands of accelerators with minimal latency and maximum reliability. Tomorrow, deploying a frontier model at scale should be as straightforward as spinning up a serverless database. The complexity doesn't disappear as it gets absorbed into the infrastructure you're building.\n\n\n\n\nSKILLS AND QUALIFICATIONS\n\nMinimum qualifications:\n\n - Bachelor's degree or equivalent experience in computer science, engineering, or similar.\n\n - Strong systems programming skills in Rust, Go, or C++.\n\n - Experience designing and building high-performance distributed systems at scale.\n\n - Understanding of network protocols and high-performance I/O.\n\n - Ability to debug complex distributed systems issues.\n\nPreferred qualifications:\n\n - Experience with ML serving infrastructure and disaggregated inference architecture.\n\n - Familiarity with GPU programming models and memory hierarchies.\n\n - Knowledge of GPU interconnects (NVLink, InfiniBand, RoCE) and their performance characteristics.\n\n - Track record of improving system reliability and performance at scale.\n\nBonus points if you have:\n\n - Prior experience in supporting large‑scale model training or inference environments.\n \n \n\n\nLOGISTICS\n\n - Location: This role is based in Singapore.\n\n - Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is S$200,000 to S$400,000 annually + equity.\n\n - Visa sponsorship: We sponsor visas on a case-by-case basis.\n\n - Benefits: Inferact offers a generous benefits package, including medical, dental, and vision coverage."},{"id":"b96d4f8b-257b-4a82-8975-57e7fd885d87","title":"Member of Technical Staff, AMD GPU Performance Engineering","department":"Research & Engineering","team":"Research & Engineering","employmentType":"FullTime","location":"Singapore","secondaryLocations":[],"publishedAt":"2026-06-25T21:49:33.825+00:00","isListed":true,"isRemote":false,"workplaceType":"OnSite","address":{"postalAddress":{"addressRegion":"Singapore","addressCountry":"Singapore","addressLocality":"Singapore"}},"jobUrl":"https://jobs.ashbyhq.com/inferact/b96d4f8b-257b-4a82-8975-57e7fd885d87","applyUrl":"https://jobs.ashbyhq.com/inferact/b96d4f8b-257b-4a82-8975-57e7fd885d87/application","descriptionHtml":"

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.

About the Role

We're looking for an AMD GPU performance engineer to make vLLM a first-class inference engine across the AMD accelerator ecosystem. You'll build and optimize AMD GPU backends, kernels, runtime paths, and benchmarking infrastructure using ROCm, HIP, Triton, CK, AITER, and related tooling so vLLM can deliver frontier inference performance on AMD GPUs.

You'll work at the boundary of inference systems, kernels, compilers, and hardware architecture, improving performance-critical paths such as attention, GEMM, sampling, KV cache, and communication-heavy operations. Your work will help make AMD GPU support in vLLM usable, fast, benchmarked, and maintainable.

Skills and Qualifications

Minimum qualifications:

Preferred qualifications:

Bonus points if you have:

Logistics

","descriptionPlain":"Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.\n\nAbout the Role\n\nWe're looking for an AMD GPU performance engineer to make vLLM a first-class inference engine across the AMD accelerator ecosystem. You'll build and optimize AMD GPU backends, kernels, runtime paths, and benchmarking infrastructure using ROCm, HIP, Triton, CK, AITER, and related tooling so vLLM can deliver frontier inference performance on AMD GPUs.\n\nYou'll work at the boundary of inference systems, kernels, compilers, and hardware architecture, improving performance-critical paths such as attention, GEMM, sampling, KV cache, and communication-heavy operations. Your work will help make AMD GPU support in vLLM usable, fast, benchmarked, and maintainable.\n\n\n\nSkills and Qualifications\n\nMinimum qualifications:\n\n - Bachelor's degree or equivalent experience in computer science, engineering, systems, machine learning, or similar.\n\n - Hands-on experience optimizing AMD GPU workloads using ROCm, HIP, Triton, CK, AITER, or similar AMD ecosystem tools.\n\n - Deep understanding of AMD GPU execution, memory behavior, toolchains, kernel performance, and backend-specific performance constraints.\n\n - Experience optimizing ML kernels or inference paths such as attention, GEMM, sampling, KV cache, fused kernels, or communication-heavy runtime paths.\n\n - Strong performance profiling and benchmarking skills, with the ability to use measurements, hardware counters, correctness tests, and reproducible benchmarks to guide optimization work.\n\nPreferred qualifications:\n\n - Experience with vLLM, SGLang, TensorRT-LLM, ROCm-based serving, or other LLM inference systems.\n\n - Familiarity with batching, KV cache, decoding, serving tradeoffs, and backend performance constraints in production inference systems.\n\n - Experience with compiler and kernel technologies such as Triton, MLIR, LLVM, CK, AITER, HIP, or other kernel DSLs and backend libraries.\n\n - Knowledge of quantization methods such as INT8, FP8, mixed precision, or AMD hardware-specific numeric formats, including accuracy and performance tradeoffs.\n\nBonus points if you have:\n\n - Contributed to vLLM, ROCm, HIP, Triton, CK, AITER, PyTorch, compiler projects, or other open-source ML infrastructure.\n\n - Built AMD GPU benchmarking infrastructure or automated performance regression detection for accelerator workloads.\n\n - Worked directly with AMD, accelerator platform teams, or early-access programs to ship backend, compiler, or inference performance improvements.\n\nLogistics\n\n - Location: This role is based in Singapore.\n\n - Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is S$200,000 to S$400,000 annually + equity.\n\n - Visa sponsorship: We sponsor visas on a case-by-case basis.\n\n - Benefits: Inferact offers a generous benefits package, including medical, dental, and vision coverage."},{"id":"76d942c3-fbb1-463d-ad79-0d9bfdcc37e9","title":"Member of Technical Staff, TPU Performance Engineering","department":"Research & Engineering","team":"Research & Engineering","employmentType":"FullTime","location":"Singapore","secondaryLocations":[],"publishedAt":"2026-06-26T03:57:59.648+00:00","isListed":true,"isRemote":false,"workplaceType":"OnSite","address":{"postalAddress":{"addressRegion":"Singapore","addressCountry":"Singapore","addressLocality":"Singapore"}},"jobUrl":"https://jobs.ashbyhq.com/inferact/76d942c3-fbb1-463d-ad79-0d9bfdcc37e9","applyUrl":"https://jobs.ashbyhq.com/inferact/76d942c3-fbb1-463d-ad79-0d9bfdcc37e9/application","descriptionHtml":"

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.

About the Role

We're looking for a TPU performance engineer to make vLLM a first-class inference engine on Google TPUs. You'll build and optimize TPU backends, compiler integrations, runtime paths, and benchmarking infrastructure using JAX, XLA, Pallas, and related tooling so vLLM can deliver frontier inference performance on TPU hardware.

You'll work at the boundary of inference systems, kernels, compilers, and hardware architecture, improving production-relevant model serving on TPU with clear correctness, latency, and throughput benchmarks. Your work will help make TPU support in vLLM usable, fast, benchmarked, and maintainable.

Skills and Qualifications

Minimum qualifications:

Preferred qualifications:

Bonus points if you have:

Logistics

","descriptionPlain":"Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.\n\nAbout the Role\n\nWe're looking for a TPU performance engineer to make vLLM a first-class inference engine on Google TPUs. You'll build and optimize TPU backends, compiler integrations, runtime paths, and benchmarking infrastructure using JAX, XLA, Pallas, and related tooling so vLLM can deliver frontier inference performance on TPU hardware.\n\nYou'll work at the boundary of inference systems, kernels, compilers, and hardware architecture, improving production-relevant model serving on TPU with clear correctness, latency, and throughput benchmarks. Your work will help make TPU support in vLLM usable, fast, benchmarked, and maintainable.\n\n\n\nSkills and Qualifications\n\nMinimum qualifications:\n\n - Bachelor's degree or equivalent experience in computer science, engineering, systems, machine learning, or similar.\n\n - Hands-on experience building or optimizing TPU workloads using JAX, XLA, Pallas, or related compiler and runtime tooling.\n\n - Deep understanding of TPU execution, memory behavior, compilation, and performance constraints for ML workloads.\n\n - Experience optimizing ML kernels or inference paths such as attention, GEMM, sampling, KV cache, fused kernels, or backend runtime paths.\n\n - Strong performance profiling and benchmarking skills, with the ability to use measurements, compiler artifacts, correctness tests, and reproducible benchmarks to guide optimization work.\n\nPreferred qualifications:\n\n - Experience with vLLM, SGLang, TensorRT-LLM, XLA-based serving, or other LLM inference systems.\n\n - Familiarity with batching, KV cache, decoding, serving tradeoffs, and backend performance constraints in production inference systems.\n\n - Experience with compiler technologies such as XLA, MLIR, LLVM, Pallas, or other kernel DSLs, including lowering, fusion, and backend code generation.\n\n - Knowledge of quantization methods such as INT8, FP8, mixed precision, or TPU-specific numeric formats, including accuracy and performance tradeoffs.\n\nBonus points if you have:\n\n - Contributed to vLLM, JAX/XLA, Pallas, PyTorch/XLA, compiler projects, or other open-source ML infrastructure.\n\n - Built TPU benchmarking infrastructure or automated performance regression detection for accelerator workloads.\n\n - Worked directly with Google TPU ecosystem stakeholders, accelerator platform teams, or early-access programs to ship backend, compiler, or inference performance improvements.\n\nLogistics\n\n - Location: This role is based in Singapore.\n\n - Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is S$200,000 to S$400,000 annually + equity.\n\n - Visa sponsorship: We sponsor visas on a case-by-case basis.\n\n - Benefits: Inferact offers a generous benefits package, including medical, dental, and vision coverage."},{"id":"1d87c80a-0e58-4745-a51f-8bd48bc73ca8","title":"Member of Technical Staff, AMD GPU Performance Engineering","department":"Research & Engineering","team":"Research & Engineering","employmentType":"FullTime","location":"San Francisco","secondaryLocations":[],"publishedAt":"2026-06-26T04:01:04.022+00:00","isListed":true,"isRemote":false,"workplaceType":"OnSite","address":{"postalAddress":{"addressRegion":"California","addressCountry":"United States","addressLocality":"San Francisco"}},"jobUrl":"https://jobs.ashbyhq.com/inferact/1d87c80a-0e58-4745-a51f-8bd48bc73ca8","applyUrl":"https://jobs.ashbyhq.com/inferact/1d87c80a-0e58-4745-a51f-8bd48bc73ca8/application","descriptionHtml":"

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.

About the Role

We're looking for an AMD GPU performance engineer to make vLLM a first-class inference engine across the AMD accelerator ecosystem. You'll build and optimize AMD GPU backends, kernels, runtime paths, and benchmarking infrastructure using ROCm, HIP, Triton, CK, AITER, and related tooling so vLLM can deliver frontier inference performance on AMD GPUs.

You'll work at the boundary of inference systems, kernels, compilers, and hardware architecture, improving performance-critical paths such as attention, GEMM, sampling, KV cache, and communication-heavy operations. Your work will help make AMD GPU support in vLLM usable, fast, benchmarked, and maintainable.

 

Skills and Qualifications

Minimum qualifications:

Preferred qualifications:

Bonus points if you have:

Logistics

","descriptionPlain":"Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.\n\nAbout the Role\n\nWe're looking for an AMD GPU performance engineer to make vLLM a first-class inference engine across the AMD accelerator ecosystem. You'll build and optimize AMD GPU backends, kernels, runtime paths, and benchmarking infrastructure using ROCm, HIP, Triton, CK, AITER, and related tooling so vLLM can deliver frontier inference performance on AMD GPUs.\n\nYou'll work at the boundary of inference systems, kernels, compilers, and hardware architecture, improving performance-critical paths such as attention, GEMM, sampling, KV cache, and communication-heavy operations. Your work will help make AMD GPU support in vLLM usable, fast, benchmarked, and maintainable.\n\n \n\nSkills and Qualifications\n\nMinimum qualifications:\n\n - Bachelor's degree or equivalent experience in computer science, engineering, systems, machine learning, or similar.\n\n - Hands-on experience optimizing AMD GPU workloads using ROCm, HIP, Triton, CK, AITER, or similar AMD ecosystem tools.\n\n - Deep understanding of AMD GPU execution, memory behavior, toolchains, kernel performance, and backend-specific performance constraints.\n\n - Experience optimizing ML kernels or inference paths such as attention, GEMM, sampling, KV cache, fused kernels, or communication-heavy runtime paths.\n\n - Strong performance profiling and benchmarking skills, with the ability to use measurements, hardware counters, correctness tests, and reproducible benchmarks to guide optimization work.\n\nPreferred qualifications:\n\n - Experience with vLLM, SGLang, TensorRT-LLM, ROCm-based serving, or other LLM inference systems.\n\n - Familiarity with batching, KV cache, decoding, serving tradeoffs, and backend performance constraints in production inference systems.\n\n - Experience with compiler and kernel technologies such as Triton, MLIR, LLVM, CK, AITER, HIP, or other kernel DSLs and backend libraries.\n\n - Knowledge of quantization methods such as INT8, FP8, mixed precision, or AMD hardware-specific numeric formats, including accuracy and performance tradeoffs.\n\nBonus points if you have:\n\n - Contributed to vLLM, ROCm, HIP, Triton, CK, AITER, PyTorch, compiler projects, or other open-source ML infrastructure.\n\n - Built AMD GPU benchmarking infrastructure or automated performance regression detection for accelerator workloads.\n\n - Worked directly with AMD, accelerator platform teams, or early-access programs to ship backend, compiler, or inference performance improvements.\n\nLogistics\n\n - Location: This role is based in San Francisco, California. Will consider remote in the US for exceptional candidates.\n\n - Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.\n\n - Visa sponsorship: We sponsor visas on a case-by-case basis.\n\n - Benefits: Inferact offers generous health, dental, and vision benefits as well as 401(k) company match."},{"id":"2eaf94cb-47f9-4aec-a293-c6052fef3511","title":"Product Marketing Manager","department":"GTM/Marketing","team":"GTM/Marketing","employmentType":"FullTime","location":"San Francisco","secondaryLocations":[],"publishedAt":"2026-07-23T00:28:56.667+00:00","isListed":true,"isRemote":false,"workplaceType":"OnSite","address":{"postalAddress":{"addressRegion":"California","addressCountry":"United States","addressLocality":"San Francisco"}},"jobUrl":"https://jobs.ashbyhq.com/inferact/2eaf94cb-47f9-4aec-a293-c6052fef3511","applyUrl":"https://jobs.ashbyhq.com/inferact/2eaf94cb-47f9-4aec-a293-c6052fef3511/application","descriptionHtml":"

Overview

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.

About the Role

We’re looking for a Product Marketing Manager to accelerate vLLM’s position as the leading open-source inference engine and publicize Inferact’s role in driving its development. This role sits at the intersection of product marketing, events, partner marketing, community, and growth. You’ll help turn a high volume of conference and ecosystem demand into a proactive GTM motion that builds brand awareness, community engagement, and qualified momentum around vLLM and Inferact

You’ll own end-to-end execution for conferences, partner events, hosted events, landing pages, speaker coordination, booth needs, collateral, follow-up workflows, and the many details that make technical events successful. You’ll develop and strengthen relationships between Inferact and major corporate and ecosystem partners, while shaping the messaging and presence of both vLLM and Inferact, while developing . You’ll work closely with product, engineering, leadership, partners, and design collaborators. This is a hands-on 0-to-1 role for someone with strong taste, high urgency, and an AI-native way of working.

Skills and Qualifications

Minimum qualifications:

Preferred qualifications:

Bonus points if you have:

Logistics

","descriptionPlain":"Overview\n\nInferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.\n\nAbout the Role\n\nWe’re looking for a Product Marketing Manager to accelerate vLLM’s position as the leading open-source inference engine and publicize Inferact’s role in driving its development. This role sits at the intersection of product marketing, events, partner marketing, community, and growth. You’ll help turn a high volume of conference and ecosystem demand into a proactive GTM motion that builds brand awareness, community engagement, and qualified momentum around vLLM and Inferact\n\n\n\nYou’ll own end-to-end execution for conferences, partner events, hosted events, landing pages, speaker coordination, booth needs, collateral, follow-up workflows, and the many details that make technical events successful. You’ll develop and strengthen relationships between Inferact and major corporate and ecosystem partners, while shaping the messaging and presence of both vLLM and Inferact, while developing . You’ll work closely with product, engineering, leadership, partners, and design collaborators. This is a hands-on 0-to-1 role for someone with strong taste, high urgency, and an AI-native way of working.\n\nSkills and Qualifications\n\nMinimum qualifications:\n\n - Bachelor's degree or equivalent experience in marketing, product, computer science, business, communications, or a related field.\n\n - Strong product marketing or GTM fundamentals, including positioning, messaging, audience definition, narrative development, and practical execution.\n\n - AI-native operating style, with regular use of modern AI tools to accelerate writing, research, planning, landing pages, workflows, asset drafts, and execution.\n\n - Experience owning events, conferences, partner activations, launches, or community programs end-to-end, including logistics, timelines, vendors, speakers, assets, landing pages, and follow-up.\n\n - Strong taste and judgment around brand, tone of voice, event experience, marketing quality, and developer-facing credibility.\n\n - Ability to operate autonomously in a startup environment, proactively recommend what should happen next, and execute without needing every step prescribed.\n\n - Ability to work credibly with product, engineering, founders, partners, design agencies, and technical communities.\n\nPreferred qualifications:\n\n - Experience marketing to developers, infrastructure engineers, ML engineers, technical founders, open-source communities, or AI / developer tools audiences.\n\n - Experience with AI infrastructure, developer tools, model serving, cloud platforms, open-source infrastructure, or other technical products.\n\n - Experience running partner marketing or co-marketing programs with ecosystem partners such as cloud providers, hardware companies, infrastructure platforms, or developer communities.\n\n - Ability to personally create or coordinate landing pages, event pages, simple websites, forms, briefs, and campaign assets using AI, no-code tools, or lightweight technical workflows.\n\n - Experience building 0-to-1 marketing programs, event playbooks, launch motions, or community programs in a startup or high-ambiguity environment.\n\nBonus points if you have:\n\n - Built an event or community strategy that materially improved brand awareness, partner momentum, developer mindshare, or qualified pipeline.\n\n - Worked in or around AI infrastructure, open-source infrastructure, developer tools, data infrastructure, cloud, GPU, or technical SaaS ecosystems.\n\n - Created high-quality technical marketing assets, launch narratives, event experiences, partner campaigns, or developer-facing content that you can walk through in detail.\n\n - Partnered with design agencies or external creative teams while maintaining speed, quality, and brand consistency.\n\n - Helped a startup move from reactive marketing execution to a proactive 6-12 month GTM, events, or community plan.\n\nLogistics\n\n - Location: This role is based in San Francisco, California. Will consider remote in the US for exceptional candidates.\n\n - Visa sponsorship: We sponsor visas on a case-by-case basis.\n\n - Benefits: Inferact offers generous health, dental, and vision benefits as well as 401(k) company match."},{"id":"ad992ead-2a9a-4694-8fca-0504354548cd","title":"Member of Technical Staff, Site Reliability Engineer","department":"Research & Engineering","team":"Research & Engineering","employmentType":"FullTime","location":"San Francisco","secondaryLocations":[],"publishedAt":"2026-08-20T18:25:22.552+00:00","isListed":true,"isRemote":false,"workplaceType":"OnSite","address":{"postalAddress":{"addressRegion":"California","addressCountry":"United States","addressLocality":"San Francisco"}},"jobUrl":"https://jobs.ashbyhq.com/inferact/ad992ead-2a9a-4694-8fca-0504354548cd","applyUrl":"https://jobs.ashbyhq.com/inferact/ad992ead-2a9a-4694-8fca-0504354548cd/application","descriptionHtml":"

Overview

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.

About the Role

We're looking for a Site Reliability Engineer to help make vLLM-powered inference systems reliable, observable, and operationally simple at production scale. This role is for someone who thinks about failure before launch, designs systems that are easier to operate, and knows how to turn incidents into durable improvements rather than one-off fixes.

You'll work across engineering and infrastructure to define SLOs, improve monitoring and alerting, strengthen incident response, drive post-mortems, and reduce operational risk before it reaches users. Your work will directly impact the reliability, availability, and production readiness of the systems powering AI inference at scale.

Skills and Qualifications

Minimum qualifications:

Preferred qualifications:

Bonus points if you have:

Logistics

","descriptionPlain":"Overview\n\nInferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.\n\nAbout the Role\n\nWe're looking for a Site Reliability Engineer to help make vLLM-powered inference systems reliable, observable, and operationally simple at production scale. This role is for someone who thinks about failure before launch, designs systems that are easier to operate, and knows how to turn incidents into durable improvements rather than one-off fixes.\n\nYou'll work across engineering and infrastructure to define SLOs, improve monitoring and alerting, strengthen incident response, drive post-mortems, and reduce operational risk before it reaches users. Your work will directly impact the reliability, availability, and production readiness of the systems powering AI inference at scale.\n\n\n\nSkills and Qualifications\n\nMinimum qualifications:\n\n - Bachelor's degree or equivalent experience in computer science, engineering, systems, infrastructure, or similar.\n\n - Strong experience operating production systems with meaningful traffic, user impact, or infrastructure criticality.\n\n - Deep understanding of SLOs, SLIs, error budgets, alerting, incident response, and post-mortem processes.\n\n - Experience live-fighting major production incidents, including mitigation, root cause analysis, escalation, and follow-through on prevention work.\n\n - Strong Linux, networking, systems debugging, observability, and distributed systems fundamentals.\n\n - Ability to design operationally simple systems and identify likely failure modes before launch.\n\n - Strong programming or scripting ability in Python, Go, Bash, or similar for automation, tooling, and reliability improvements.\n\nPreferred qualifications:\n\n - Experience supporting ML infrastructure, inference systems, GPU workloads, Kubernetes-based platforms, or high-scale backend services.\n\n - Experience building or improving observability systems using metrics, logs, traces, dashboards, alerts, and runbooks.\n\n - Experience with Kubernetes, Docker, Terraform, cloud infrastructure, service meshes, CI/CD systems, or production deployment platforms.\n\n - Experience driving incident review culture, post-mortem processes, reliability reviews, and prevention-oriented engineering work.\n\n - Ability to partner with engineering teams to improve service design, release safety, capacity planning, and operational readiness.\n\nBonus points if you have:\n\n - Owned reliability for high-throughput, latency-sensitive, or mission-critical production systems.\n\n - Supported AI inference, model serving, GPU clusters, ML platforms, or distributed serving infrastructure.\n\n - Built automation that reduced toil, improved recovery time, or prevented repeat incidents.\n\n - Led incident response for severe outages with clear communication across engineering and leadership.\n\n - Created practical SLOs, dashboards, alerts, runbooks, or release gates that improved production reliability.\n\n\n\nLogistics\n\n - Location: This role is based in San Francisco, California. Will consider remote in the US for exceptional candidates.\n\n - Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.\n\n - Visa sponsorship: We sponsor visas on a case-by-case basis.\n\n - Benefits: We offers generous health, dental, and vision benefits as well as 401(k) company match."},{"id":"20df5968-cafb-4007-a6f6-ad854411b501","title":"Head of Engineering","department":"Research & Engineering","team":"Research & Engineering","employmentType":"FullTime","location":"San Francisco","secondaryLocations":[],"publishedAt":"2026-08-03T07:16:47.827+00:00","isListed":true,"isRemote":false,"workplaceType":"OnSite","address":{"postalAddress":{"addressRegion":"California","addressCountry":"United States","addressLocality":"San Francisco"}},"jobUrl":"https://jobs.ashbyhq.com/inferact/20df5968-cafb-4007-a6f6-ad854411b501","applyUrl":"https://jobs.ashbyhq.com/inferact/20df5968-cafb-4007-a6f6-ad854411b501/application","descriptionHtml":"

Overview

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.

About the Role

We're looking for a Head of Engineering to build and lead the organization developing the systems that power vLLM and Inferact. This role requires an engineering leader with genuine technical credibility at the inference layer—someone who understands GPU and accelerator performance, inference runtimes, ML systems optimization, and hardware-software co-design deeply enough to earn the trust of exceptional staff-level engineers.

You'll partner closely with the founders to scale a senior-heavy, highly specialized engineering team while preserving the technical rigor, speed, and ownership that made vLLM successful. You'll recruit and develop rare ML systems talent, translate ambitious research and infrastructure work into a focused execution plan, strengthen how teams operate, and help Inferact deliver reliable, high-performance inference across models, hardware, and deployment environments.

Skills and Qualifications

Minimum qualifications:

Preferred qualifications:

Bonus points if you have:

Logistics

","descriptionPlain":"Overview\n\nInferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.\n\nAbout the Role\n\nWe're looking for a Head of Engineering to build and lead the organization developing the systems that power vLLM and Inferact. This role requires an engineering leader with genuine technical credibility at the inference layer—someone who understands GPU and accelerator performance, inference runtimes, ML systems optimization, and hardware-software co-design deeply enough to earn the trust of exceptional staff-level engineers.\n\nYou'll partner closely with the founders to scale a senior-heavy, highly specialized engineering team while preserving the technical rigor, speed, and ownership that made vLLM successful. You'll recruit and develop rare ML systems talent, translate ambitious research and infrastructure work into a focused execution plan, strengthen how teams operate, and help Inferact deliver reliable, high-performance inference across models, hardware, and deployment environments.\n\n\n\nSkills and Qualifications\n\nMinimum qualifications:\n\n - Bachelor's degree or equivalent experience in computer science, engineering, machine learning, systems, or a related field.\n\n - Engineering leadership experience building and scaling highly specialized teams in LLM inference, ML systems, GPU or accelerator software, distributed systems, or closely related infrastructure.\n\n - Deep technical credibility at the inference layer, including hands-on understanding of inference runtimes, GPU or accelerator optimization, kernels, memory and communication bottlenecks, and hardware-software tradeoffs.\n\n - Ability to distinguish core inference-engine work from the routing, orchestration, and application layers above it, with opinions grounded in direct technical experience.\n\n - A strong record of recruiting, assessing, and retaining senior engineers, staff-level ICs, PhDs, and research-adjacent engineers in a production engineering environment.\n\n - Experience translating technically ambitious work into clear priorities, accountable ownership, execution plans, and durable engineering operating mechanisms.\n\n - Ability to remain close enough to the work to identify risks, pattern-match on difficult technical problems, and unblock teams without becoming a bottleneck or displacing technical ownership.\n\nPreferred qualifications:\n\n - Experience leading teams responsible for LLM serving, vLLM, SGLang, model execution, inference performance, GPU kernels, compiler or runtime systems, or distributed AI infrastructure.\n\n - Experience scaling a small, senior-heavy engineering organization where the relevant talent market is narrow and technical quality matters more than headcount growth.\n\n - Experience integrating research-oriented or PhD talent into production teams, including setting expectations, structuring work, and building effective collaboration with product-focused engineers.\n\n - Strong judgment across organizational design, hiring, performance management, technical planning, execution cadence, and cross-functional decision-making.\n\n - Ability to represent the engineering organization credibly with open-source contributors, hardware partners, cloud providers, customers, candidates, and investors.\n\nBonus points if you have:\n\n - Built or led engineering teams working directly on GPU or accelerator-level inference performance, ML compilers, kernels, runtimes, or hardware-software co-design.\n\n - Contributed to or led teams around open-source ML systems projects such as vLLM, SGLang, PyTorch, Ray, Triton, XLA, ROCm, or related infrastructure.\n\n - Scaled an engineering organization through an inflection point while preserving high technical standards, fast iteration, and direct ownership.\n\n - Recruited successfully from a global, highly competitive ML systems talent pool and built relationships with technical communities beyond traditional candidate pipelines.\n\n - Led engineering in an early-stage AI infrastructure, developer infrastructure, distributed systems, or open-source company.\n\n\n\nLogistics\n\n - Location: This role is based in San Francisco, California. Will consider relocation for exceptional candidates.\n\n - Compensation: Compensation will be determined based on background, skills, and experience. Offer will include a highly competitive base and meaningful equity.\n\n - Visa sponsorship: We sponsor visas on a case-by-case basis.\n\n - Benefits: Inferact offers generous health, dental, and vision benefits as well as 401(k) company match."},{"id":"f73d3f45-cdea-41a8-8144-acf504aa4fda","title":"Founding Product Designer","department":"Product","team":"Product","employmentType":"FullTime","location":"San Francisco","secondaryLocations":[],"publishedAt":"2026-08-08T00:17:50.028+00:00","isListed":true,"isRemote":false,"workplaceType":"OnSite","address":{"postalAddress":{"addressRegion":"California","addressCountry":"United States","addressLocality":"San Francisco"}},"jobUrl":"https://jobs.ashbyhq.com/inferact/f73d3f45-cdea-41a8-8144-acf504aa4fda","applyUrl":"https://jobs.ashbyhq.com/inferact/f73d3f45-cdea-41a8-8144-acf504aa4fda/application","descriptionHtml":"

Overview

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.

About the Role

We’re looking for a Founding Product Designer to give Inferact and vLLM a distinct visual identity, sharp brand presence, and highly intuitive developer product experiences. This is a 0-to-1, full-stack design role for a creator with high aesthetic taste, deep technical curiosity, and an AI-native way of executing.

You will bridge the gap between technical infrastructure and visual craft. In your first 90 days, you’ll establish our visual language across high-visibility social assets, model launch campaigns, blog graphics, and partner announcements. As we scale, you’ll transition into shaping our web surfaces, brand identity, internal tooling, and core product interfaces. You’ll work directly with founders, engineering leads, product, and marketing to build design systems that scale across open-source communities, developer tools, and enterprise platforms.

30-60-90 Day Milestones

Skills and Qualifications

Minimum qualifications:

Preferred qualifications:

Bonus points if you have:

Logistics

Location: This role is based in San Francisco, California. Will consider remote in the US for exceptional candidates.

Compensation: Compensation is to be determined based on background, skills, and experience. Offers include equity.

Visa sponsorship: We sponsor visas on a case-by-case basis.

Benefits: Inferact offers generous health, dental, and vision benefits as well as 401(k) company match.

","descriptionPlain":"Overview\n\nInferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.\n\n\n\nAbout the Role\n\nWe’re looking for a Founding Product Designer to give Inferact and vLLM a distinct visual identity, sharp brand presence, and highly intuitive developer product experiences. This is a 0-to-1, full-stack design role for a creator with high aesthetic taste, deep technical curiosity, and an AI-native way of executing.\n\nYou will bridge the gap between technical infrastructure and visual craft. In your first 90 days, you’ll establish our visual language across high-visibility social assets, model launch campaigns, blog graphics, and partner announcements. As we scale, you’ll transition into shaping our web surfaces, brand identity, internal tooling, and core product interfaces. You’ll work directly with founders, engineering leads, product, and marketing to build design systems that scale across open-source communities, developer tools, and enterprise platforms.\n\n\n\n30-60-90 Day Milestones\n\n - 30 Days: Own and ship high-impact design assets for model launches, technical content, and ecosystem news. You’ll onboard to the inference ecosystem by establishing a baseline visual language for Inferact and vLLM.\n\n - 60 Days: Expand into initial UI/UX designs for our production tooling and developer dashboards, while continuing to increase your design footprint across our core website, interactive developer pages, partner announcement assets, and i.\n\n - 90 Days: Own the core design for our enterprise product: focus on core product design—refining UI/UX for developer tools, optimizing developer workflows, and establishing a unified design system from brand assets to product components.\n\n\n\nSkills and Qualifications\n\nMinimum qualifications:\n\n - Bachelor's degree or equivalent practical experience in Product Design, Human-Computer Interaction (HCI), Computer Science, Interactive Media, or a related field.\n\n - A strong portfolio or tech-forward personal website demonstrating exceptional craft, strong typography, layout, visual systems, and interactive UI/UX thinking.\n\n - AI-native operating style, using modern AI image, design, and code generators to explore concepts rapidly and ship production-ready assets.\n\n - High technical taste and fluency, with a track record of designing for developers, technical products, or complex systems.\n\n - End-to-end execution capability—able to take a rough idea or technical blog post and independently produce polished visuals, landing pages, or product mockups without needing prescribed steps.\n\n - Ability to collaborate credibly with technical founders, engineers, and product managers in a fast-paced, high-urgency startup environment.\n\nPreferred qualifications:\n\n - Experience designing for developer tools, AI infrastructure, cloud platforms, open-source communities, or technical SaaS products.\n\n - Strong background in visual design, brand systems, illustration, or graphic assets alongside digital product design (UI/UX).\n\n - Proficiency with modern web design and front-end prototyping tools (Figma, Framer, Webflow, React/Tailwind code prototypes).\n\n - Experience designing graphics and collateral for major developer conferences, community events, and launch campaigns.\n\nBonus points if you have:\n\n - Motion design skills (After Effects, Rive, Lottie) or lightweight 3D design experience (Blender, Cinema 4D, Spline).\n\n - Built and maintained a personal tech-forward website or creative engineering projects showcasing custom interaction design.\n\n - Created 0-to-1 visual identities or design systems for a high-growth startup or prominent open-source project.\n\n\n\nLogistics\n\nLocation: This role is based in San Francisco, California. Will consider remote in the US for exceptional candidates.\n\nCompensation: Compensation is to be determined based on background, skills, and experience. Offers include equity.\n\nVisa sponsorship: We sponsor visas on a case-by-case basis.\n\nBenefits: Inferact offers generous health, dental, and vision benefits as well as 401(k) company match.\n\n"},{"id":"595cbea0-7099-4416-a87c-efcc2876e654","title":"Member of Technical Staff, Cluster Administration","department":"Research & Engineering","team":"Research & Engineering","employmentType":"FullTime","location":"San Francisco","secondaryLocations":[],"publishedAt":"2026-08-21T16:13:04.388+00:00","isListed":true,"isRemote":false,"workplaceType":"OnSite","address":{"postalAddress":{"addressRegion":"California","addressCountry":"United States","addressLocality":"San Francisco"}},"jobUrl":"https://jobs.ashbyhq.com/inferact/595cbea0-7099-4416-a87c-efcc2876e654","applyUrl":"https://jobs.ashbyhq.com/inferact/595cbea0-7099-4416-a87c-efcc2876e654/application","descriptionHtml":"

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware—a position that took years to build.

About the Role

We're looking for a hands-on cluster administration engineer to own and operate the high-performance GPU compute infrastructure that keeps Inferact engineering productive. Inferact runs on expensive, high-performance GPU and HPC clusters across neo-cloud and dedicated compute providers. Your job is to make sure that infrastructure is healthy, available, observable, and usable around the clock.

You'll take ownership of cluster health, GPU availability, monitoring, alerting, scheduling, access, diagnostics, and incident response across the systems our engineers rely on every day. You'll work closely with engineering leadership and infrastructure owners to standardize how we provision, operate, debug, and scale compute across providers. Your work will directly impact how fast Inferact can build, test, and improve the systems powering vLLM.

Skills and Qualifications

Minimum qualifications:

Preferred qualifications:

Bonus points if you have:

Logistics

","descriptionPlain":"Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware—a position that took years to build.\n\n\n\nAbout the Role\n\nWe're looking for a hands-on cluster administration engineer to own and operate the high-performance GPU compute infrastructure that keeps Inferact engineering productive. Inferact runs on expensive, high-performance GPU and HPC clusters across neo-cloud and dedicated compute providers. Your job is to make sure that infrastructure is healthy, available, observable, and usable around the clock.\n\nYou'll take ownership of cluster health, GPU availability, monitoring, alerting, scheduling, access, diagnostics, and incident response across the systems our engineers rely on every day. You'll work closely with engineering leadership and infrastructure owners to standardize how we provision, operate, debug, and scale compute across providers. Your work will directly impact how fast Inferact can build, test, and improve the systems powering vLLM.\n\n\n\nSkills and Qualifications\n\nMinimum qualifications:\n\n - Bachelor's degree or equivalent experience in computer science, engineering, systems administration, or similar.\n\n - Hands-on experience administering large compute clusters, HPC environments, university or research clusters, supercomputing systems, or production GPU clusters.\n\n - Strong Linux systems administration fundamentals across networking, processes, storage, package management, shell scripting, logs, access control, and system debugging.\n\n - Experience operating GPU servers, including driver management, GPU health monitoring, node failures, memory errors, scheduler issues, and hardware diagnostics.\n\n - Experience with cluster scheduling and resource allocation using SLURM, Kubernetes, or equivalent tooling.\n\n - Ability to own urgent infrastructure incidents end-to-end when compute issues are blocking engineering teams.\n\n - Ability to automate operational workflows using Bash, Python, Ansible, Terraform, Helm, or similar tooling.\n\nPreferred qualifications:\n\n - Experience operating GPU compute across providers such as Lambda, CoreWeave, Crusoe, Nebius, Together, Fireworks, RunPod, or similar environments.\n\n - Experience improving cluster utilization, reducing idle or unavailable GPU capacity, and debugging scheduling or resource contention issues.\n\n - Familiarity with high-performance GPU networking such as InfiniBand, RoCE, NVLink / NVSwitch, RDMA, NCCL, or equivalent systems.\n\n - Experience with storage for HPC or ML workloads, including NFS, Lustre, Ceph, distributed filesystems, or other high-throughput storage systems.\n\n - Experience managing secure access, identity, permissions, SSH, VPNs, bastion hosts, secrets, and basic infrastructure security hygiene.\n\n - Background in research computing, scientific computing, ML infrastructure, SRE, platform engineering, or infrastructure operations for engineering-heavy teams.\n\nBonus points if you have:\n\n - Managed GPU or HPC infrastructure in a university lab, national lab, research institution, AI infrastructure company, hedge fund, HFT firm, or large-scale ML platform team.\n\n - Built monitoring, alerting, runbooks, health checks, or remediation workflows that materially reduced operational toil or incident resolution time.\n\n - Operated Kubernetes clusters for ML or GPU workloads at meaningful scale.\n\n - Standardized provisioning, diagnostics, monitoring, and operating patterns across multiple compute providers.\n\n - Carried real operational responsibility for infrastructure used by many engineers or researchers.\n\n\n\nLogistics\n\n - Location: This role is based in San Francisco, California. Will consider remote in the US for exceptional candidates.\n\n - Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.\n\n - Visa sponsorship: We sponsor visas on a case-by-case basis.\n\n - Benefits: Inferact offers generous health, dental, and vision benefits as well as 401(k) company match."},{"id":"910f4dbb-29ba-4105-a3e7-076fba07123a","title":"HR / People Lead","department":"Human Resources","team":"Human Resources","employmentType":"FullTime","location":"San Francisco","secondaryLocations":[],"publishedAt":"2026-09-09T19:58:43.769+00:00","isListed":true,"isRemote":false,"workplaceType":"OnSite","address":{"postalAddress":{"addressRegion":"California","addressCountry":"United States","addressLocality":"San Francisco"}},"jobUrl":"https://jobs.ashbyhq.com/inferact/910f4dbb-29ba-4105-a3e7-076fba07123a","applyUrl":"https://jobs.ashbyhq.com/inferact/910f4dbb-29ba-4105-a3e7-076fba07123a/application","descriptionHtml":"

Overview

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.

About the Role

We are looking for an experienced, high-ownership HR / People Lead to own, design, and scale our People Operations and HR infrastructure as we expand. In this key role, you will bridge the critical handoff from talent acquisition to employee onboarding and serve as the core operational owner for the entire employee lifecycle.

You will work directly with leadership, finance, and engineering teams to manage HR operations end-to-end—spanning onboarding workflows, employee relations, payroll, benefits, compliance, and immigration/visa filings. The ideal candidate brings a blend of strategic HR experience from scaling tech companies and the scrappy, hands-on execution needed in an early-stage startup environment.

Skills and Qualifications

Minimum qualifications:

Preferred qualifications:

Logistics

","descriptionPlain":"Overview\n\nInferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.\n\n\n\nAbout the Role\n\nWe are looking for an experienced, high-ownership HR / People Lead to own, design, and scale our People Operations and HR infrastructure as we expand. In this key role, you will bridge the critical handoff from talent acquisition to employee onboarding and serve as the core operational owner for the entire employee lifecycle.\n\nYou will work directly with leadership, finance, and engineering teams to manage HR operations end-to-end—spanning onboarding workflows, employee relations, payroll, benefits, compliance, and immigration/visa filings. The ideal candidate brings a blend of strategic HR experience from scaling tech companies and the scrappy, hands-on execution needed in an early-stage startup environment.\n\n\n\nSkills and Qualifications\n\nMinimum qualifications:\n\n - Bachelor's degree or equivalent practical experience.\n\n - 4+ years of experience as an HR Business Partner, People Specialist, or People Operations lead in fast-paced tech environments.\n\n - Proven hands-on experience managing core HR infrastructure: onboarding, payroll, benefits, employee relations, and compliance.\n\n - Direct experience utilizing modern startup tools and platforms, including PEO, payroll, benefits, ATS, etc.\n\n - Experience managing immigration and visa workflows in coordination with external legal counsel.\n\n - Exceptional interpersonal and communication skills, with a track record of building trust across founders, employees, and external vendors.\n\n\n\nPreferred qualifications:\n\n - Experience spanning both high-growth startups and established tech companies.\n\n - Demonstrated ability to operate as a sole IC or strategic lead in a fast-moving 0-to-1 environment.\n\n - Familiarity with cross-border HR logistics, international PEO structures, or working with APAC HR partners.\n\n\n\nLogistics\n\n - Location: San Francisco, CA (On-site / Hybrid).\n\n - Compensation: $180,000 – $250,000 base salary based on experience and background, plus meaningful startup equity.\n\n - Benefits: Generous health, dental, and vision insurance coverage, alongside a 401(k) company match."},{"id":"894029d4-5381-4fd7-9b28-eed06307f796","title":"Head of Legal","department":"Legal","team":"Legal","employmentType":"FullTime","location":"San Francisco","secondaryLocations":[],"publishedAt":"2026-09-09T21:00:13.840+00:00","isListed":true,"isRemote":false,"workplaceType":"OnSite","address":{"postalAddress":{"addressRegion":"California","addressCountry":"United States","addressLocality":"San Francisco"}},"jobUrl":"https://jobs.ashbyhq.com/inferact/894029d4-5381-4fd7-9b28-eed06307f796","applyUrl":"https://jobs.ashbyhq.com/inferact/894029d4-5381-4fd7-9b28-eed06307f796/application","descriptionHtml":"

Overview

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.

About the Role

We're looking for our first in-house legal hire to give Inferact greater leverage across a growing set of legal workstreams. This is a hands-on, high-impact role for an experienced technology lawyer who can personally execute on commercial agreements while serving as the primary point person for all legal matters at the company.

This hire will provide immediate in-house legal capacity and, in short order, build out and lead Inferact's legal function. You will own key matters internally, coordinate complex workstreams with outside counsel, maintain cross-functional alignment across the organization, and help the company navigate frontier questions in open source LLM inference space. You will work closely with leadership, product, engineering, GTM, and operations teams.

Skills and Qualifications

Minimum qualifications:

Preferred qualifications:

Bonus points if you have:

Logistics

Location: San Francisco, CA (On-site / Hybrid). Will consider remote in the US for exceptional candidates.

Compensation: Compensation determined based on background, skills, and experience. Offers include competitive base salary & meaningful equity.

Visa sponsorship: We sponsor visas on a case-by-case basis.

Benefits: Inferact offers generous health, dental, and vision benefits as well as 401(k) company match.

","descriptionPlain":"Overview\n\nInferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.\n\n\n\nAbout the Role\n\nWe're looking for our first in-house legal hire to give Inferact greater leverage across a growing set of legal workstreams. This is a hands-on, high-impact role for an experienced technology lawyer who can personally execute on commercial agreements while serving as the primary point person for all legal matters at the company.\n\nThis hire will provide immediate in-house legal capacity and, in short order, build out and lead Inferact's legal function. You will own key matters internally, coordinate complex workstreams with outside counsel, maintain cross-functional alignment across the organization, and help the company navigate frontier questions in open source LLM inference space. You will work closely with leadership, product, engineering, GTM, and operations teams.\n\n\n\nSkills and Qualifications\n\nMinimum qualifications:\n\n - J.D. or equivalent, with active bar admission and good standing in at least one relevant jurisdiction.\n\n - 6+ years of legal experience at a technology company, infrastructure or AI business, or a top law firm advising sophisticated tech clients. Prior GC experience is not required.\n\n - Experience building or scaling a legal function from 0 to 1, or operating as the sole or first lawyer supporting a fast-growing company.\n\n - Strong commercial contracting experience: able to independently draft, negotiate, and close enterprise customer, cloud and compute vendor, partnership, and technology licensing agreements.\n\n - Experience managing outside counsel efficiently, with clear internal ownership and a bias toward keeping matters moving.\n\n - Hands-on execution mindset—comfortable personally drafting, reviewing, and negotiating in a technically complex, fast-moving startup.\n\n - Strong business judgment, with the ability to translate legal complexity into clear, practical tradeoffs for leadership and engineers.\n\nPreferred qualifications:\n\n - Experience at or advising companies in AI inference, model serving, GPU cloud, or developer infrastructure.\n\n - Working fluency in open-source software licensing and the legal mechanics of running a company alongside a community-driven open-source project.\n\n - Familiarity with the regulatory issues facing AI infrastructure companies, including data privacy, export controls on compute and models, and emerging AI-specific regulation.\n\n - Experience negotiating with hyperscalers, hardware vendors, and enterprise customers with demanding security, compliance, and data-handling requirements.\n\n - Track record of building lightweight self-service resources, templates, and playbooks so product, engineering, and go-to-market teams can move without waiting on legal.\n\nBonus points if you have:\n\n - Served as first or sole in-house counsel, or as a senior lawyer with broad ownership, at an AI infrastructure company\n\n - Deep experience with open-source business models (open core, managed offerings, enterprise licensing, etc.) or companies that steward a major open-source project.\n\n - Handled the full generalist range at an early-stage startup: financings, equity and employment, immigration, and corporate governance.\n\n - Built strong working relationships with top law firms while keeping accountability and execution speed in-house.\n\n\n\nLogistics\n\nLocation: San Francisco, CA (On-site / Hybrid). Will consider remote in the US for exceptional candidates.\n\nCompensation: Compensation determined based on background, skills, and experience. Offers include competitive base salary & meaningful equity.\n\nVisa sponsorship: We sponsor visas on a case-by-case basis.\n\nBenefits: Inferact offers generous health, dental, and vision benefits as well as 401(k) company match.\n\n"},{"id":"c0f52cce-53fc-478c-894e-1ee601148cc6","title":"Member of Technical Staff, Cloud Orchestration (Remote)","department":"Research & Engineering","team":"Research & Engineering","employmentType":"FullTime","location":"Remote","secondaryLocations":[],"publishedAt":"2026-08-25T05:48:30.953+00:00","isListed":true,"isRemote":true,"workplaceType":"Remote","address":{"postalAddress":{"addressCountry":"US"}},"jobUrl":"https://jobs.ashbyhq.com/inferact/c0f52cce-53fc-478c-894e-1ee601148cc6","applyUrl":"https://jobs.ashbyhq.com/inferact/c0f52cce-53fc-478c-894e-1ee601148cc6/application","descriptionHtml":"

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware—a position that took years to build.

About the Role

We're looking for an cloud orchestration engineer to build the operational backbone that keeps vLLM running reliably at massive scale. You'll design the systems for cluster management, deployment automation, and production monitoring that enable teams worldwide to serve AI models without friction. You'll ensure that vLLM deployments are observable, debuggable, and recoverable, turning operational complexity into infrastructure that just works.

Skills and Qualifications

Minimum qualifications:

Preferred qualifications:

Bonus points if you have:

Logistics

","descriptionPlain":"Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware—a position that took years to build.\n\n\n\n\nABOUT THE ROLE\n\nWe're looking for an cloud orchestration engineer to build the operational backbone that keeps vLLM running reliably at massive scale. You'll design the systems for cluster management, deployment automation, and production monitoring that enable teams worldwide to serve AI models without friction. You'll ensure that vLLM deployments are observable, debuggable, and recoverable, turning operational complexity into infrastructure that just works.\n\n\n\n\nSKILLS AND QUALIFICATIONS\n\nMinimum qualifications:\n\n - Bachelor's degree or equivalent experience in computer science, engineering, or similar.\n\n - Strong experience with Kubernetes and container orchestration at scale.\n\n - Experience designing and implementing custom Kubernetes operators.\n\n - Proficiency in Python/Rust/Go and infrastructure-as-code tools (Terraform, Helm, etc).\n\n - Experience managing GPU clusters and debugging hardware issues.\n\n - Ability to work across cloud platforms (AWS, GCP, Azure) and on-premise infrastructure.\n\nPreferred qualifications:\n\n - Experience with ML-specific orchestration tools (Ray, Slurm).\n\n - Knowledge of GPU scheduling, multi-tenancy, and resource optimization.\n\n - Familiarity with vLLM deployment patterns and configuration.\n\n - Track record of improving operational reliability for ML systems.\n\nBonus points if you have:\n\n - Experience deploying inference systems on large-scale GPU (1,000+) clusters.\n \n \n\n\nLOGISTICS\n\n - Location: Fully remote, worldwide. We're timezone-flexible but expect regular overlap with Pacific Time for critical syncs.\n\n - Compensation: We offer competitive compensations (salary + equity) compared to the local market conditions.\n\n - Visa sponsorship: We sponsor visas on a case-by-case basis.\n\n - Benefits: Inferact offers competitive benefits appropriate to your location, including health coverage where applicable."},{"id":"ef7198da-ad0a-4c02-963f-6250a15e3534","title":"Member of Technical Staff, Inference","department":"Research & Engineering","team":"Research & Engineering","employmentType":"FullTime","location":"Remote","secondaryLocations":[],"publishedAt":"2026-09-04T07:59:03.927+00:00","isListed":true,"isRemote":true,"workplaceType":"Remote","address":{"postalAddress":{"addressCountry":"US"}},"jobUrl":"https://jobs.ashbyhq.com/inferact/ef7198da-ad0a-4c02-963f-6250a15e3534","applyUrl":"https://jobs.ashbyhq.com/inferact/ef7198da-ad0a-4c02-963f-6250a15e3534/application","descriptionHtml":"

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware—a position that took years to build.

About the Role

We're looking for an inference runtime engineer to push the boundaries of what's possible in LLM and diffusion model serving. Models grow larger. Architectures shift: mixture-of-experts, multimodal, agentic. Every breakthrough demands innovations on the inference engine itself. You'll work at the core of vLLM, optimizing how models execute across diverse hardware and architectures. Your work will directly impact how the world runs AI inference.

Skills and Qualifications

Minimum qualifications:

Preferred qualifications:

Bonus points if you have:

Logistics

","descriptionPlain":"Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware—a position that took years to build.\n\n\n\n\nABOUT THE ROLE\n\nWe're looking for an inference runtime engineer to push the boundaries of what's possible in LLM and diffusion model serving. Models grow larger. Architectures shift: mixture-of-experts, multimodal, agentic. Every breakthrough demands innovations on the inference engine itself. You'll work at the core of vLLM, optimizing how models execute across diverse hardware and architectures. Your work will directly impact how the world runs AI inference. \n\n\n\n\nSKILLS AND QUALIFICATIONS\n\nMinimum qualifications:\n\n - Bachelor's degree or equivalent experience in computer science, engineering, or similar.\n\n - Deep understanding of transformer architectures and their variants.\n\n - Strong programming skills in Python with experience in PyTorch internals.\n\n - Experience with LLM inference systems (vLLM, TensorRT-LLM, SGLang, TGI).\n\n - Ability to read and implement model architectures and inference techniques from research papers.\n\n - Demonstrate the ability to contribute performant and maintainable code and debug in complex ML codebases.\n\nPreferred qualifications:\n\n - Deep understanding of KV-cache memory management, prefix caching, and hybrid model serving.\n\n - Familiarity with RL frameworks and algorithms for LLMs.\n\n - Experience with multimodal inference (audio/image/video/text).\n\n - Contributions to open-source ML or system infrastructure projects.\n\nBonus points if you have:\n\n - Implemented core features in vLLM or other inference engine projects.\n\n - Contributed to vLLM integrations (verl, OpenRLHF, Unsloth, LlamaFactory, etc).\n\n - Written widely-shared technical blogs or side projects on vLLM or LLM inference.\n \n \n\n\nLOGISTICS\n\n - Location: Fully remote, worldwide. We're timezone-flexible but expect regular overlap with Pacific Time for critical syncs.\n\n - Compensation: We offer competitive compensations (salary + equity) compared to the local market conditions.\n\n - Visa sponsorship: We sponsor visas on a case-by-case basis.\n\n - Benefits: Inferact offers competitive benefits appropriate to your location, including health coverage where applicable."}],"apiVersion":"1"}