Staff Applied AI Inference Engineer
crusoeLakewood, CO
yesterday
Occupations
Software DevelopersComputer Systems Engineers/ArchitectsData ScientistsIndustries
Custom Computer Programming ServicesComputer Systems Design ServicesComputing Infrastructure Providers, Data Processing, Web Hosting, and Related ServicesAbout the role
Overview
In this role you will optimize large language model inference in production, reducing time and cost while maintaining reliability. You will own the end-to-end inference stack, from profiling to deployment, and work closely with customer teams to tailor deployments. The role blends core systems work with customer-facing and product efforts to deliver measurable performance gains. You will operate in a fast-paced, hands-on environment, delivering production-ready optimizations that scale across models and workloads.
Compensation / Benefits Competitive compensation and equity packages Comprehensive health, dental & vision insurance 401(k) Retirement plan with company match Paid time off & holidays Professional development & tuition reimbursement Parental leave
Responsibilities Own the end-to-end inference stack, profiling time and cost and applying optimizations in production Design and optimize serving architectures (prefill, decode disaggregation, request routing)Dig into the serving stack from frameworks like vLLM and SGLang to CUDA kernels to identify performance issues Adapt optimization methods to a range of ML models with emphasis on large language models Profile and tune deployments against latency, throughput, and cost targets under real traffic Tailor deployments to customer models and constraints from PoC to live production service Build and support software/product features around the inference stack in production using Python (preferred)Experiment quickly, turn fuzzy goals into concrete specs, run proofs of concept, and ship tested results Own delivery end-to-end from experiment to production optimization and draft product requirements with teams Make sound trade-offs and reduce unnecessary complexity Show ownership and accountability in work and team culture
Key requirements Bachelor's in Computer Science, Engineering, Mathematics, or related field Production code experience in Python or C++ (Python preferred)Familiarity with LLM optimization for high throughput/low latency inference Experience with LLM serving frameworks (vLLM, SGLang) and kernel-level performance analysis Strong understanding of GPU architecture Hands-on experience with large language models and AI/ML pipelines Strong communication skills for explaining technical topics to customers and teammatesproblem-solving mindsetstrong communicationcustomer-facing collaboration PythonC++CUDA
Matching similar jobs
JOB OVERVIEW
Experience level
Lead
Location
Lakewood, CO
Occupation
Software Developers
Industry
Custom Computer Programming Services
Posted
yesterday
Tired of running searches?
Rank the roles you'd take once, and matches like these arrive on their own.
CREATE PROFILE