LLM Inference Engineer: Scale & Optimize Production
comunidade metodistaPalo Alto, CA
LLM Inference Engineer: Scale & Optimize Production
comunidade metodistaPalo Alto, CA
2 days ago
Occupations
Computer Systems Engineers/ArchitectsSoftware DevelopersComputer and Information Research ScientistsIndustries
Computing Infrastructure Providers, Data Processing, Web Hosting, and Related ServicesSoftware PublishersElectronic Computer ManufacturingAbout the role
Hippocratic AI is seeking an experienced LLM Inference Engineer to optimize its LLM serving infrastructure in Palo Alto, CA. You will design multi-node architectures, implement multi-LoRA serving, and apply quantization techniques to reduce footprint while maintaining quality.
The role emphasizes practical production deployment, benchmarking, and latency-focused optimizations across GPUs, with a strong emphasis on Python, C++, and CUDA.
Matching similar jobs
JOB OVERVIEW
Experience level
Lead
Location
Palo Alto, CA
Occupation
Computer Systems Engineers/Architects
Industry
Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services
Posted
2 days ago
Tired of running searches?
Rank the roles you'd take once, and matches like these arrive on their own.
CREATE PROFILE