Staff/Senior Software Engineer, Document Intelligence & Data Infrastructure
AI summary of the role
Senior engineering role owning the acquisition and processing infrastructure for Helios' Proxi platform, turning global public-sector and open-source data into searchable intelligence.
What you’ll do
- Expand Proxi's ingestion and document-processing platform from source acquisition through downstream publication.
- Construct crawling and connector infrastructure for discovering, acquiring, and maintaining international data sources.
- Scale distributed processing for document conversion, extraction, transcription, enrichment, replay, and backfills.
- Establish common data contracts that normalize heterogeneous multilingual sources while preserving provenance.
What you’ll bring
- Built and operated large-scale acquisition or document-processing systems in production.
- Expertise in web crawlers/connectors for continuously changing international sources.
- Expertise in fault-tolerant pipelines with reliable scheduling, recovery, replay, and backfill.
- Expertise in document conversion, OCR, layout analysis, and structural extraction across complex file formats.
Technologies
web crawlers · OCR · document conversion · distributed processing · real-time transcription · speaker diarization · GPU inference · GovCloud · air-gapped environments · multilingual processing
Source and classification
Internal deployment & tooling · Evidence for this classification:
We bring together frontier AI, deep public-sector expertise, and forward-deployed execution. Our team includes leaders and builders from the White House, U.S. Department of State, Datadog, and Microsoft. We are backed by leading institutional investors and trusted by organizations working on high-stakes problems across government and industry. MISSION Build and operate the acquisition and processing infrastructure that turns global public-sector and open-source data into reliable, searchable intelligence. This role owns the path from source discovery and collection through normalization, extraction, enrichment, and publication. You will expand our platform to include new countries, languages, institutions, and source formats covering many additional areas of open source intelligence to advance Proxi’s natural capabilities. We are looking for a senior engineer who has built and
More from the job description
Staff/Senior Software Engineer, Document Intelligence & Data Infrastructure New York City | Full-time | On-site in SoHo, five days per week Reports to the CTO and works directly with the founding team ABOUT HELIOS Helios is building a new kind of company to solve America’s hardest problems, starting with the government interaction layer. Government shapes every consequential market, but the infrastructure connecting public institutions and private organizations remains fragmented, manual, and difficult to navigate. Helios is rebuilding that layer. Our core platform, Proxi, gives organizations the intelligence they need to understand what government is doing, why it matters, and what to do next. From that foundation, we design and deploy secure, mission-specific systems for government agencies, enterprises, and institutions operating in complex and highly regulated environments. We bring together frontier AI, deep public-sector expertise, and forward-deployed execution. Our team includes leaders and builders from the White House, U.S. Department of State, Datadog, and Microsoft. We are backed by leading institutional investors and trusted by organizations working on high-stakes problems across government and industry. MISSION Build and operate the acquisition and processing infrastructure that turns global public-sector and open-source data into reliable, searchable int [... source excerpt omitted ...] rce intelligence to advance Proxi’s natural capabilities. We are looking for a senior engineer who has built and operated large-scale acquisition or document-processing systems in production. Expected areas of expertise: Web crawlers and connectors for continuously changing international sources. Fault-tolerant pipelines with reliable scheduling, recovery, replay, and backfill capabilities. Processing complex documents, structured files, images, audio, and video across languages. Stable schemas, source lineage, correction propagation. Operating secure, observable, scalable, and cost-efficient processing systems across cloud and restricted environments. Document conversion, OC [... source excerpt omitted ...] , versioning, correction propagation, and reproducible reprocessing. Secure handling of untrusted content with rigorous quality evaluation, observability, and cost controls. KEY RESPONSIBILITIES Expand Proxi’s ingestion and document-processing platform from source acquisition through downstream publication. Construct crawling and connector infrastructure required to discover, acquire, and continuously maintain international data sources. Scale the distributed processing environment that supports document conversion, extraction, transcription, enrichment, replay, and backfills. Establish common data contracts that normalize heterogeneous and multilingual sources without discardin
Employer postings · Data from · Sources