Skip navigation
#213365

Sr. AI Infrastructure Engineer

Warren, MI OR Austin, TX
Date:

Overview

Placement Type:

Temporary

Salary:

$80-85 Hourly

Start Date:

Oct 12, 2026

Join a pioneering team at the forefront of operational excellence, where your expertise will directly empower thousands of leaders across a global manufacturing network. This is a unique opportunity to shape the future of intelligent systems, driving efficiency and innovation that resonates throughout a vast enterprise, and make a tangible difference in daily operations.

About Our Client

Our client, a global leader in the manufacturing sector, is dedicated to revolutionizing its operations through cutting-edge technology and intelligent solutions. With a mission to enhance productivity and streamline processes across its extensive global footprint, they are building advanced platforms that leverage artificial intelligence to deliver unprecedented insights and operational efficiency. Partnering with Aquent, they seek visionary talent to drive this transformative journey.

About the Role

We are seeking an exceptionally talented and passionate engineer to join a dynamic team building the next generation of AI-powered operational platforms. In this pivotal role, you will be instrumental in developing, deploying, and managing the infrastructure that underpins sophisticated AI products. You’ll contribute to a critical operational platform designed to provide a single, connected experience for operational leaders, directly impacting efficiency across hundreds of global facilities. This is your chance to apply your deep technical skills to complex, real-world challenges, eliminating administrative burdens and fostering a more efficient, data-driven environment.

What You’ll Do

  • Lead the software engineering production of advanced artificial intelligence products, ensuring robust and scalable solutions.
  • Support comprehensive development operations (DevOps) and release management for AI products and a broader product portfolio.
  • Design, implement, and optimize distributed systems, focusing on reliability and high performance.
  • Conduct rigorous testing of features and capabilities to uphold product quality and system credibility.
  • Engage in daily stand-up meetings with a highly collaborative team, managing tasks through modern project management tools like Jira.
  • Utilize cutting-edge AI coding tools to accelerate development and delivery of innovative solutions.
  • Contribute to the development of critical modules, such as staffing solutions, that directly impact operational efficiency on the plant floor.
  • Collaborate extensively with diverse teams, including engineers and managers, to drive project success.

What You’ll Bring (Must-Have Qualifications)

  • Bachelor’s degree in a technical field such as Computer Science, Computer Engineering, or a related discipline.
  • 8-10 years of overall experience in software engineering, with at least 6 years focused on Infrastructure, Site Reliability Engineering (SRE), or Core Platform Engineering, particularly with distributed systems.
  • Demonstrated experience with AI-driven development, AI engineering, and coding, especially within the last few years of employment.
  • Expertise in software engineering principles, including a strong developer mindset with proficiency in Python and React.
  • Advanced proficiency with Kubernetes and Helm, including a deep understanding of the K8s control plane, custom controllers/CRDs, and experience tuning stateful sets or long-lived connection routing.
  • Production-level Python experience (FastAPI/Uvicorn), showcasing expert knowledge of Python’s memory management, asyncio event loops, and tuning ASGI servers for high-concurrency gateway flows.
  • Proven experience in advanced load testing execution using tools like k6 or Locust to test stateful systems, handle dynamic session tokens, and mock asynchronous backend workers.
  • Hands-on automation experience with CI/CD pipelines and container hardening, specifically with GitHub Actions targeting Azure Container Registry (ACR) or similar enterprise registries, and Azure Kubernetes Service (AKS).
  • Deep conceptual understanding of LLM (Large Language Model) infrastructure, including tool-calling, agent memory architectures, state graphs, and the operational differences between token-streaming versus unary API calls.
  • Exceptional communication and collaboration skills, with a proven ability to work effectively with cross-functional teams.
  • A strong ability to leverage AI tools to accelerate development processes and deliver high-quality solutions efficiently.

Bonus Points (Nice-to-Have Qualifications)

  • Production experience with Service Mesh technologies such as Istio, Linkerd, or Cilium for advanced A2A traffic splitting, mutual TLS (mTLS), and zero-trust enforcement.
  • Experience with agent sandboxing using secure container runtimes (e.g., gVisor, Kata Containers) for executing unverified user-generated tool code or agent scripts.
  • Practical experience with Chaos Engineering, including injecting failures into asynchronous message queues or mocking LLM API timeouts to validate graph recovery safety.

Why This Role Matters

This is more than just a job; it’s an opportunity to be at the forefront of a major new initiative with significant business impact. You’ll join a highly senior and collaborative team of 12 engineers, contributing to a project that will be rolled out to hundreds of operational leaders globally. The environment is complex and fast-paced, offering continuous learning and growth. There’s a strong likelihood for extension into next year, with potential opportunities for direct hire, making this an exceptional path for career advancement.

Interview Process

  • First Round: A 30-minute virtual interview via Teams with the hiring manager, focusing on behavioral questions (3-5 questions).
  • Second Round: A 30-minute virtual technical interview via Teams conducted by technical leads.

Other Requirements

Unless expressly authorized in the interview/skill set validation instructions, candidates may not use AI tools or other third-party assistance during an interview or assessment to generate or provide substantive answers to interview questions. Candidates should answer organically in their own words, drawing on their own knowledge, skills, and experience. Failure to adhere to these expectations may result in the candidate being removed from consideration.

Note: This is not meant to address use of assistive technologies (e.g., captioning, screen readers, speech-to-text, text-to-speech, etc.) where approved as reasonable accommodations.