Remote job
Senior Platform Engineer
Job details
About this role
Role overview
A senior platform engineer role focused on scaling the infrastructure that powers a domain-specific AI platform serving the architecture, engineering, and construction industry. The platform combines embedding models, document parsing, and autonomous agents that reason over real-world project data, and is already deployed across enterprise customers on three continents. The position is a senior individual contributor with broad ownership and real architectural influence, where the infrastructure decisions made define how reliably the system runs and scales in the field.
Responsibilities
- Orchestrate rollouts of agents and services across many customer cloud environments, including deployment strategies, per-customer configuration, automated health checks, and proactive monitoring that catches problems before customers do - Keep model inference fast, reliable, and cost-effective at scale, covering serving infrastructure, GPU workloads, and ongoing performance work as volume grows - Manage core infrastructure spanning Kubernetes, multi-account AWS, CI/CD, observability (traces, metrics, logs, alerting, SLOs), disaster recovery, and cost management - Strengthen security posture through access controls, secrets management, network security, image scanning, dependency auditing, and compliance efforts such as SOC 2 as enterprise requirements demand - Define, provision, and evolve all infrastructure through code, designing modules, managing state, and reasoning carefully about blast radius
Requirements
- 5+ years in infrastructure, DevOps, or SRE roles running cloud infrastructure in production - Strong Kubernetes experience deploying workloads, debugging real issues, and working with operators and controllers - Solid infrastructure-as-code skills covering module design, state management, and reasoning about blast radius - Strong software engineering fundamentals with the ability to write and review production code in Python and/or TypeScript - Linux systems and networking fundamentals - CI/CD pipeline design and maintenance experience - A proactive orientation with genuine comfort owning a wide surface area
Nice to have
- Terraform experience - Observability platforms such as Datadog or OpenTelemetry, including dashboards and trace, metric, and log pipelines - PostgreSQL operations such as performance tuning and replica management - ML or AI infrastructure experience with inference services, GPU workloads, model serving, or evaluation pipelines - Multi-tenant deployment patterns or per-customer isolation - Experience building sandboxed execution environments or automated reliability systems
Benefits and work setup
- Competitive base salary and performance-based compensation - Equity participation - Medical, dental, and vision coverage - Flexible paid time off - Hybrid NYC model preferred, remote considered for the right candidate