← Back to jobs

Remote job

Senior Inference Engineer

AI Engineer Full-time Permanent United States

Job details

$190K – $220K • Offers Equity • Offers Bonus Salary
United States Eligibility
Senior Experience
Full-time Employment

About this role

Role overview This is the founding engineering hire for an inference platform, partnering directly with the CTO to build the production inference layer from scratch. The role owns the full pipeline that turns a customer's query into a served response at scale, then leads the technical direction of the inference platform going forward. It is an engineering role, not a research position, with deep hands-on responsibility for serving infrastructure and optimization.

Responsibilities - Deploy large language models into production across one or more GPU-backed machines, owning the end-to-end query-to-response pipeline. - Stand up and operate serving infrastructure using modern LLM serving frameworks. - Apply quantization, batching, caching, and routing techniques to control latency and cost as traffic grows. - Write production-grade infrastructure code in Python or Golang, not just configuration. - Build the first iteration alongside the CTO, then lead the inference roadmap and partner with Product on future direction. - Communicate technical concepts clearly to both engineers and non-technical stakeholders.

Requirements - Hands-on experience deploying and serving LLMs in production using vLLM, SGLang, or TensorRT-LLM, ideally at a company built around inference at scale. - Practical experience with inference optimization techniques including quantization, batching, caching, and routing. - Strong production engineering skills in Python or Golang, with real shipped code behind you. - Excellent written and verbal communication skills for explaining technical tradeoffs to varied audiences.

Nice to have - Familiarity with containerized environments such as Docker and Kubernetes. - Generative AI experience with ML frameworks like PyTorch and Transformers. - Working knowledge of the GPU stack, including CUDA, NCCL, drivers, and related libraries. - Understanding of model architectures and fine-tuning approaches. - Experience with NVIDIA Dynamo.

Benefits and work setup - Competitive compensation package that includes equity. - Platinum-level health, dental, vision, and life insurance coverage for employees and eligible dependents, with details varying by country. - Flexible working schedule focused on outcomes rather than rigid hours. - Workplace flexibility and support for adjusting the work environment as life changes. - Remote-first culture with global team distribution.

Skills detected in the listing

PythonGoStakeholder ManagementDockerKubernetesLLM
Detected Sep 11, 2026
Last verified Sep 11, 2026

Hidden Jobs Access

Unlock application links

Read the full job details for free. An active Hidden Jobs Access subscription is required to open the original application link.

Weekly

FREE $6.99/week after trial
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching
  • Cancel anytime before day 7

Monthly

$35.99 $17.99 /month
  • 35% cheaper than weekly
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching

Lifetime

$99.99 $49.99 /forever
  • One-time payment
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching
Hidden Jobs gives subscribers direct access to original application links
Offer ends in 00:00:00 Your profile-fit rate expires at midnight