Staff Applied AI Inference Engineer

crusoeLakewood, CO

yesterday

Occupations

Software DevelopersComputer Systems Engineers/ArchitectsData Scientists

Industries

Custom Computer Programming ServicesComputer Systems Design ServicesComputing Infrastructure Providers, Data Processing, Web Hosting, and Related Services
APPLY NOW

About the role

Overview In this role you will optimize large language model inference in production, reducing time and cost while maintaining reliability. You will own the end-to-end inference stack, from profiling to deployment, and work closely with customer teams to tailor deployments. The role blends core systems work with customer-facing and product efforts to deliver measurable performance gains. You will operate in a fast-paced, hands-on environment, delivering production-ready optimizations that scale across models and workloads. Compensation / Benefits Competitive compensation and equity packages Comprehensive health, dental & vision insurance 401(k) Retirement plan with company match Paid time off & holidays Professional development & tuition reimbursement Parental leave Responsibilities Own the end-to-end inference stack, profiling time and cost and applying optimizations in production Design and optimize serving architectures (prefill, decode disaggregation, request routing)Dig into the serving stack from frameworks like vLLM and SGLang to CUDA kernels to identify performance issues Adapt optimization methods to a range of ML models with emphasis on large language models Profile and tune deployments against latency, throughput, and cost targets under real traffic Tailor deployments to customer models and constraints from PoC to live production service Build and support software/product features around the inference stack in production using Python (preferred)Experiment quickly, turn fuzzy goals into concrete specs, run proofs of concept, and ship tested results Own delivery end-to-end from experiment to production optimization and draft product requirements with teams Make sound trade-offs and reduce unnecessary complexity Show ownership and accountability in work and team culture Key requirements Bachelor's in Computer Science, Engineering, Mathematics, or related field Production code experience in Python or C++ (Python preferred)Familiarity with LLM optimization for high throughput/low latency inference Experience with LLM serving frameworks (vLLM, SGLang) and kernel-level performance analysis Strong understanding of GPU architecture Hands-on experience with large language models and AI/ML pipelines Strong communication skills for explaining technical topics to customers and teammatesproblem-solving mindsetstrong communicationcustomer-facing collaboration PythonC++CUDA

Matching similar jobs

JOB OVERVIEW

Experience level

Lead

Location

Lakewood, CO

Occupation

Software Developers

Industry

Custom Computer Programming Services

Posted

yesterday

Tired of running searches?

Rank the roles you'd take once, and matches like these arrive on their own.

CREATE PROFILE