Remote job
Staff Software Engineer, Data Platform
Job details
About this role
Role overview A senior individual contributor role on the Datastore team at a brain-data platform serving clinical research and digital diagnostics. The position focuses on the design and scaling of backend data infrastructure, including PostgreSQL data models, Kafka-driven event pipelines, and a GraphQL API that makes scientific and clinical data usable across the organization. The role is 100% remote within the United States, with optional in-person collaboration at offices in Boston, New York City, and Paris.
Responsibilities - Lead design and architecture for complex systems within the Datastore, weighing trade-offs against engineering principles and production data. - Partner with product managers, scientists, clinicians, and other engineering teams to translate requirements into new integrations, services, and platform capabilities. - Debug and profile ambiguous problems that cross system boundaries, leaving code, tests, and tooling in better shape than found. - Draft RFCs to drive feature discovery and technical alignment across the team. - Own pipelines that deliver dataset snapshots into the warehouse and serve as a central source for downstream consumers. - For one of the two openings, focus deeply on PostgreSQL data modeling, performance, and operational excellence; for the other, focus broadly across event pipelines, GraphQL, and third-party integrations.
Requirements - Strong track record designing and scaling complex backend systems, not just individual components. - Deep PostgreSQL fluency, including query planning, index selection, EXPLAIN (ANALYZE), schema design, and migration under load. - Hands-on experience with Kafka or similar event-driven pipelines, including idempotency, replay, and consumer lag handling. - Proficiency with GraphQL and TypeScript or JavaScript on Node.js for candidates focused on the API surface. - Comfort with row-level security, schema-as-API decisions, and reasoning about access controls in shared clinical datasets. - Experience or strong interest in applying LLM-assisted or agentic coding tools with appropriate guardrails.
Nice to have - Operating PostgreSQL in production, including replication, backup and restore, connection pooling, and vacuum tuning. - Designing schemas that model time correctly for re-scored and corrected clinical results. - Teaching and documenting database practices for the rest of the team.
Benefits and work setup - 100% remote within the U.S. with robust asynchronous work practices; optional hubs in Boston, NYC, and Paris. - US-based salary range of $170,000 – $190,000, plus equity, PTO, and additional benefits.