Software Engineer
AI summary of the role
Generalist software engineer on the core platform team at Radical AI, building backend services, data pipelines, internal platforms, and observability for an autonomous materials R&D lab.
What you’ll do
- Build and own backend services, agent tooling, data and orchestration pipelines, and internal platforms for the autonomous lab platform.
- Design and implement lab backend features: experiment definitions, sample path-planning, and long-running durable task execution.
- Develop data backend: database schemas, migrations, ETL pipelines, object storage and partition design.
- Implement observability and reliability: structured logs, metrics, tracing, production debugging (OpenTelemetry, Prometheus/Grafana).
What you’ll bring
- 5–6+ years of production software engineering experience, able to design, build, and ship end-to-end.
- Fluency in Go and/or Python.
- Deep comfort with distributed systems: timeouts, retries, idempotency, partial failure, dead-letter queues, safe rollback.
- Experience with concurrent and asynchronous programming, including event loops, cancellation semantics, bounded queues, task orchestration under failure.
Technologies
Go · Python · TypeScript · Rust · Kubernetes · AWS · OpenTelemetry · Prometheus · Grafana · MongoDB · gRPC · Ray
About Radical AI
Autonomous materials R&D platform combining AI, robotics, and closed-loop experimentation to discover novel inorganic materials in months instead of decades.
Seed
Source and classification
Internal deployment & tooling · Evidence for this classification:
Radical AI is replacing an R&D process that currently takes 10+ years and $100 million to produce a single discovery. Our self-driving lab platform combines AI with autonomous robotics to run experiments, analyze results, and iterate—continuously, without human bottlenecks. For industries like aerospace, automotive, defense, energy, manufacturing, semiconductors, and space, that means breakthroughs in weeks instead of years. The Role This is a generalist software engineering role on the team building the core platform that powers our autonomous lab. Depending on where you plug in, you might own backend services, agent tooling, data and orchestration pipelines, internal platforms, or the infrastructure that keeps our systems observable and reliable in production. What's non-negotiable: strong systems thinking, production instincts, and the ability to contribute across the stack. We're
More from the job description
Radical AI is replacing an R&D process that currently takes 10+ years and $100 million to produce a single discovery. Our self-driving lab platform combines AI with autonomous robotics to run experiments, analyze results, and iterate—continuously, without human bottlenecks. For industries like aerospace, automotive, defense, energy, manufacturing, semiconductors, and space, that means breakthroughs in weeks instead of years. The Role This is a generalist software engineering role on the team building the core platform that powers our autonomous lab. Depending on where you plug in, you might own backend services, agent tooling, data and orchestration pipelines, internal platforms, or the infrastructure that keeps our systems observable and reliable in production. What's non-negotiable: strong systems thinking, production instincts, and the ability to contribute across the stack. We're building software that controls and monitors physical systems in the real world — the bar for correctness and reliability is high. We work across a deeply cross-disciplinary domain (robotics, ML, experimental automation), and one near-term priority involves hybrid cloud/on-prem deployments in customer environments. Engineers who've navigated those constraints will hit the ground running. What You'll Work On Lab backend: experiment definitions, sample path-planning, long running durable task e [... source excerpt omitted ...] ject storage + partition design Internal platforms: developer tooling, SDKs, shared services, service templates Observability and reliability: structured logs, metrics, tracing, production debugging (OpenTelemetry, Prometheus/Grafana) Hybrid infrastructure: cloud + on-prem, containerization, orchestration, infrastructure-as-code Agent capabilities and tooling: API integrations, code execution, scientific literature retrieval, workflow automation Scientific workflow orchestration: Bayesian optimization loops, experiment scheduling, long-running job execution, retries, idempotency Data pipelines for ingesting, transforming, and serving data to models and LLMs What We're Lookin
Employer postings · Data from · Sources