Inference Engineer: High-Throughput ML Serving
adaption labsMillbrae, CA
Inference Engineer: High-Throughput ML Serving
adaption labsMillbrae, CA
2 days ago
Occupations
Computer Systems Engineers/ArchitectsSoftware DevelopersNetwork and Computer Systems AdministratorsIndustries
Computing Infrastructure Providers, Data Processing, Web Hosting, and Related ServicesComputer Systems Design ServicesSoftware PublishersAbout the role
Adaption Labs, Inc. is seeking an ML systems engineer to own the cost and performance of the inference stack. You will optimize caching, batching, quantization, decoding, and kernel-level tuning to improve throughput and latency while preserving model quality.
You will collaborate with the serving fleet engineers and tackle real production workloads, focusing on cost-efficiency, tail latency, and reliable delivery across changing workloads and hardware. Bay Area presence required.
Matching similar jobs
JOB OVERVIEW
Experience level
Mid
Location
Millbrae, CA
Occupation
Computer Systems Engineers/Architects
Industry
Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services
Posted
2 days ago
Tired of running searches?
Rank the roles you'd take once, and matches like these arrive on their own.
CREATE PROFILE