Summary
Build and scale a production multi-agent AI platform serving thousands of internal users across multiple business units. Monthly release cadence, real users, real latency, real cost.
What You'll Own
- LLM-driven orchestrator that routes user intent across a portfolio of specialized agents — delegation, memory, response validation, capability discovery.
- Agent selection layer — hybrid retrieval (vector RAG over a capability registry) plus closed-set LLM selection with JSON-schema-constrained outputs.
- Multi-agent SDK / gateway — FastAPI service hosting many agents behind path-prefix routing, per-agent tool registries, session-scoped conversational context.
- Tool-driven agents — 15–30 tools per agent composed dynamically by an LLM; owns tool contracts, guardrails, and evaluation.
- Data API layer — parameterized endpoints between agents and databases; LLMs never touch DBs directly.
- Partner-team onboarding — versioned A2A contract, bring-your-own-agent registration, auto re-embedding.
Core AI Engineering
- Production LLM systems: RAG, tool/function-calling loops, structured outputs, hallucination guards, closed-set selection.
- Multi-agent orchestration: A2A protocols, session affinity, human-in-the-loop gating, kill switches, graceful degradation.
- Vector search + embeddings at scale (sub-second retrieval over thousands of docs).
- Evaluation & safety: PII/PHI masking, audit trails, feedback-loop instrumentation, offline + online eval.
Platform / Infrastructure
- Python 3.11+, FastAPI, async I/O, Pydantic.
- Modern LLM stacks (Gemini, GPT, Claude) and agent frameworks (LangGraph, Agent SDKs).
- Cloud (GCP or AWS): Kubernetes, object storage, workflow orchestration, Vertex/Bedrock-class services.
- Redis, MongoDB, Oracle/Postgres, SSO + RBAC.
- Observability: Prometheus, structured JSON logs, per-decision audit trails, p95 latency SLOs in seconds.
Ways of Working — Fast Turnaround, Ship-Fast
- Comfortable with short cycle times: spec → design → merged → deployed in days, not sprints. Monthly releases are the floor, not the ceiling.
- Bias to ship the smallest correct thing, verify in production, iterate. No polish before proof.
- Owns the full loop: intake → spec → design → implementation → code review → test evidence → UAT → deploy → post-release observation.
- Fluent with AI-assisted developer tooling (Claude Code, Cursor, agentic IDEs); reads and writes code with an LLM in the loop as a force multiplier.
Skill Curation & Reuse — Agentic Development Discipline
- Uses and extends the team's agentic SDLC skill library — capability intake, spec authoring, design docs, implementation plans, release-impact artifacts, deployment records.
- Curates new skills when a workflow repeats: codifies patterns (accessibility, security/STRIDE, CI/CD, data-source adapters, renderer standards) into reusable skills the whole team can invoke.
- Treats skills, prompts, and evals as first-class artifacts — versioned, reviewed, and improved like code.
- Knows when to reach for a skill vs. write ad-hoc: standard flows for standard work, creative bandwidth saved for novel problems.
You'll Thrive Here If
- You've shipped LLM agents in production (not demos) with real users, latency, and cost constraints.
- You reason about routing, tool selection, and context strategy as first-class design surfaces — not just prompt tuning.
- You own both the model layer and the platform underneath it (queues, auth, secrets, deploy, K8s).
- You move fast without breaking discipline: verification before completion, evidence before assertions.
- Regulated-domain experience (healthcare, financial services) is a plus.
Bonus
- Contributions to agent frameworks, evaluation harnesses, or open A2A protocols.
- Patent / IP work in agentic systems, RAG, or multi-agent orchestration.
- Prior experience authoring internal skill libraries, agent playbooks, or SDLC automation for AI teams.
Role Expectations
- Strong hands-on experience in designing, developing, and implementing Agentic AI solutions and frameworks at enterprise scale.
- Proven track record of delivering Agentic AI use cases in production environments, demonstrating measurable business value and outcomes.
- Ability to work independently with a high degree of ownership, accountability, and self-motivation.
- Experience leading or contributing to large-scale digital transformation initiatives for Fortune 500 and enterprise clients.
- Strong communication skills with the ability to clearly articulate solution approaches, business use cases, technical contributions, and delivered impact.
- Deep understanding of AI solution architecture, with the capability to confidently explain design decisions, technology choices, challenges, mitigations, and results.
- Demonstrated capability to conceptualize, architect, and build AI-driven solutions end-to-end, from ideation through deployment and adoption.
- Ability to collaborate effectively with business and technology stakeholders while driving innovation and delivering tangible business outcomes.