Remote job
#663 Backend AI Engineers
Job details
About this role
Role overview
A backend engineering role focused on designing and building a next-generation reliability platform for production systems. The position blends traditional distributed systems engineering with AI-assisted development to give teams a single place to understand, debug, and improve service health. It suits an engineer who enjoys rapid iteration and shipping while maintaining a very high bar for reliability, quality, and maintainability.
Responsibilities
- Design and build a centralized reliability command center that gives teams a unified view of system health, risk, and reliability across services and environments. - Create AI agents that assist with incident triage, root-cause exploration, log and trace summarization, and recommended next actions. - Develop developer-facing features and APIs that help engineers explore data, debug issues, and make better decisions. - Use AI-assisted development tools as leverage to prototype, refactor, and ship high-quality code quickly. - Own projects end-to-end, from requirements and architecture through implementation, testing, rollout, and iteration based on feedback. - Collaborate closely with partner teams such as product, infrastructure, data, and SRE to translate pain points into simple, powerful solutions.
Requirements
- 5+ years of experience in backend or full-stack engineering, with a track record of building and operating complex production systems. - Strong proficiency in Python, with experience architecting data-intensive applications and robust APIs. - Strong problem-solving and product sense, with the ability to take ambiguous requirements and rapidly iterate toward a working solution. - Hands-on experience with AI-assisted development tools such as Cursor, Claude, or similar, and enthusiasm for using them to ship features faster. - Practical experience using LLMs or AI frameworks to enhance automation and guidance with appropriate guardrails and citations.
Nice to have
- Background in observability, incident response, or internal developer platforms. - Comfort working across infrastructure and application boundaries in cloud data environments.