Skip to content
NEXTMOVEFDE careers · United States

Edge AI/Model Optimization Engineer

AI summary of the role

This role focuses on deploying, optimizing, and sustaining AI and agentic AI capabilities on edge and tactical computing platforms, particularly the X9 Spider Mission Computer, for defense missions.

No longer in the current catalog. Last included 2026-09-09. Check the employer’s posting for availability.

What you’ll do

  • Evaluate LLMs, embedding models, and inference solutions on GPU-enabled edge platforms like X9 Spider.
  • Tune runtime configurations (quantization, batching, caching, GPU memory) for edge deployment.
  • Benchmark agentic AI workflows and inference pipelines against hardware constraints.
  • Build and maintain performance and stress-testing frameworks for edge scenarios.

What you’ll bring

  • Bachelor's degree in CS, EE, CE, Data Science, AI, or related.
  • 5+ years in AI/ML deployment, model optimization, edge computing, GPU acceleration, or inference ops.
  • Experience deploying/optimizing LLMs or inference pipelines in resource-constrained environments.
  • Experience with CUDA, TensorRT, ONNX Runtime, vLLM, Ollama, or similar.

Technologies

Edge AI · LLM · Embedding models · GPU · CUDA · TensorRT · ONNX Runtime · vLLM · Ollama · Docker · Kubernetes · Python

About NextGen Federal Systems

Privately held federal IT services firm delivering software, R&D, systems engineering, and cyber/data solutions to DoD and the intelligence community.

Bootstrapped · 200–500 people

Source and classification

Internal deployment & tooling · Evidence for this classification:

NextGen is seeking a highly motivated and technically skilled Edge AI/Model Optimization Engineer to support the deployment, optimization, and sustainment of AI and agentic AI capabilities within edge and tactical computing environments. This role focuses on evaluating, tuning, benchmarking, and operationalizing Large Language Models (LLMs), embedding models, and AI inference services for constrained hardware platforms, including the X9 Spider Mission Computer architecture and other edge compute systems supporting operational missions using ReadiChat. ReadiChat is a mission-focused, agentic AI platform designed to help organizations build, deploy, govern, and scale specialized AI agents for operational workflows. It combines AI agents, workflow orchestration, grounded knowledge, testing frameworks, and enterprise controls into a single collaborative workspace. The ideal candidate will
More from the job description

NextGen is seeking a highly motivated and technically skilled Edge AI/Model Optimization Engineer to support the deployment, optimization, and sustainment of AI and agentic AI capabilities within edge and tactical computing environments. This role focuses on evaluating, tuning, benchmarking, and operationalizing Large Language Models (LLMs), embedding models, and AI inference services for constrained hardware platforms, including the X9 Spider Mission Computer architecture and other edge compute systems supporting operational missions using ReadiChat. ReadiChat is a mission-focused, agentic AI platform designed to help organizations build, deploy, govern, and scale specialized AI agents for operational workflows. It combines AI agents, workflow orchestration, grounded knowledge, testing frameworks, and enterprise controls into a single collaborative workspace. The ideal candidate will possess expertise in AI model optimization, GPU-enabled edge computing, runtime performance tuning, and operational AI deployment. This role requires close collaboration with AI engineers, systems integrators, mission stakeholders, and operational users to ensure AI-enabled capabilities remain performant, reliable, and mission-effective within disconnected, degraded, intermittent, and low-bandwidth environments. Responsibilities Evaluate candidate Large Language Models (LLMs), embedding models [... source excerpt omitted ...] ching configurations, context window sizing, cache behavior, inference scheduling, and GPU memory utilization specific to operational edge hardware environments. Collaborate with customer stakeholders to assess mission requirements and evaluate alternative edge compute platforms when operational demands exceed X9 Spider capabilities or when cost, performance, power, size, weight, or thermal tradeoffs require additional analysis. Benchmark agentic AI workflows, inference pipelines, and model-serving architectures against target hardware constraints and operational performance thresholds. Recommend model-selection, runtime, and configuration tradeoffs balancing mission effective [... source excerpt omitted ...] ment, troubleshooting, optimization, and sustainment activities for AI-enabled applications operating in edge, airborne, tactical, or disconnected operational environments. Train customer technical personnel on supported model profiles, operational constraints, runtime tuning considerations, deployment limitations, troubleshooting procedures, and platform sustainment best practices. Maintain technical documentation, benchmarking results, model validation reports, deployment procedures, optimization baselines, configuration guides, and operational support materials. Support DevSecOps and CI/CD activities associated with AI model packaging, deployment automation, runtime validat

How jobs are selected

Employer postings · Data from · Sources