About the role

Hippocratic AI is seeking an experienced LLM Inference Engineer to optimize its LLM serving infrastructure in Palo Alto, CA. You will design multi-node architectures, implement multi-LoRA serving, and apply quantization techniques to reduce footprint while maintaining quality. The role emphasizes practical production deployment, benchmarking, and latency-focused optimizations across GPUs, with a strong emphasis on Python, C++, and CUDA.

Matching similar jobs

JOB OVERVIEW

Experience level

Lead

Location

Palo Alto, CA

Occupation

Computer Systems Engineers/Architects

Industry

Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services

Posted

2 days ago

Tired of running searches?

Rank the roles you'd take once, and matches like these arrive on their own.

CREATE PROFILE