Remote job
Software Engineer - Research Data Platform
Job details
About this role
Role overview
A research-driven organization at the intersection of artificial intelligence and biology is hiring an engineer to build the data backbone behind its foundation models and biomedical discovery tools. The role focuses on designing, scaling, and operating storage and access layers that make large, complex biological datasets usable for scientists, product engineers, and AI agents alike. It is a mid-level to senior individual contributor position with remote flexibility across Europe.
Responsibilities
- Design and maintain scalable data schemas and storage layers optimized for high-throughput AI training and inference workloads - Profile, benchmark, and improve distributed storage using chunking, compression, parallelization, and other performance techniques - Build clean, typed, well-documented APIs that let researchers, engineers, and AI agents query and extend the platform programmatically - Partner closely with researchers and product engineers to scope requirements, align on priorities, and drive projects from design through delivery - Monitor and tune database queries, read and write paths, pipeline bottlenecks, and cloud infrastructure costs - Collaborate with platform engineering to implement data validation, testing, versioning, and access controls across the stack
Requirements
- Production-level proficiency in Python with a passion for clean, readable, and maintainable code - Hands-on experience with modern Python data tooling such as Pydantic, SQLAlchemy, Alembic, object storage abstractions, and FastAPI or comparable frameworks - Strong knowledge of relational database systems like PostgreSQL, including schema design, indexing strategies, and query optimization - Familiarity with distributed storage formats and engineering for very large datasets - Self-directed, curious, detail-oriented mindset suited to a dynamic, fast-paced research environment - Comfort working across disciplines and translating research needs into robust engineering decisions
Nice to have
- Exposure to biological, omics, or other life sciences data structures and ontologies - Experience designing APIs that serve both human users and AI agents from a single surface - Background in cloud cost optimization, infrastructure-as-code, or data governance frameworks
Benefits and work setup
- Remote-first arrangement with flexibility for candidates located across Europe - Competitive compensation package including equity participation - Mission-driven, multidisciplinary team working at the frontier of AI and biomedical research - Inclusive hiring practices with equal opportunity commitments throughout the recruitment process