Machine Learning Engineer (Agent & Inference)
Technology
Bitus Labs
Irvine, United States4 days agoUntil 10/7/2026
Full timeOn-site
Job description
Requirements
Must have:
- We require a bachelors or masters degree in Computer Science, Machine Learning, or a closely related discipline.
- We look for at least 3 years of industry experience in machine learning engineering or a comparable role.
- We need strong software and systems engineering skills, including experience building low-latency, dependable production services in Go, Rust, C++, or similar languages.
- We expect experience building or supporting real-time inference systems for recommendation, ranking, contextual bandits, reinforcement learning, or other adaptive ML applications.
- We require strong hands-on experience with PyTorch and the Hugging Face ecosystem.
- We need experience creating production LLM or agent applications using frameworks such as LangGraph, LlamaIndex, or equivalent.
- We expect practical experience with RAG systems, embeddings, and vector databases.
- We require experience evaluating and monitoring LLM or agent systems in production.
- We need experience deploying and tuning production machine learning or LLM systems.
- We expect a solid understanding of inference runtime behavior, resource usage, latency tuning, and production serving performance.
- We require experience with Docker and Kubernetes.
- We need experience with cloud platforms such as AWS, GCP, or Azure.
- We require fluency in Mandarin Chinese.
- Nice-to-have experience includes fine-tuning open-weight LLMs with LoRA, QLoRA, PEFT, or similar methods.
- Additional preferred experience includes familiarity with recommender systems, ranking systems, contextual bandits, or reinforcement learning algorithms.
- We value experience with custom GPU kernel work using CUDA or OpenAI Triton.
- We prefer experience with graph-level optimization and low-level inference performance tuning.
- We welcome experience with large-scale distributed training such as FSDP, DeepSpeed, or multi-GPU workloads.
- We value experience deploying models to edge environments using TFLite, CoreML, or NPU accelerators.
- Strong CI/CD knowledge and deployment workflow experience is preferred.
- Background in gaming, gaming AI, or player personalization systems is a plus.
- Experience with distributed systems, Spark, Hadoop, or large-scale data infrastructure is preferred.
Responsibilities:
- We design, build, and optimize LLM-powered agents, including planning, tool use, workflow orchestration, and multi-step reasoning.
- We architect memory systems covering short-term memory, long-term memory, context management, and session state.
- We build and improve RAG pipelines to strengthen relevance, grounding, freshness, and retrieval quality.
- We design and run vector-store infrastructure such as pgvector, Milvus, Qdrant, or Weaviate.
- We define evaluation methods for agents, prompts, and workflows.
- We improve end-to-end agent quality, latency, reliability, and operating cost.
- We build and operate production inference services that are low-latency, high-concurrency, and highly reliable.
- We serve online-learning models, including contextual bandits and reinforcement learning policies, with real-time inference and live parameter or weight updates.
- We deploy and optimize AI inference systems for latency, throughput, reliability, and resource efficiency.
- We analyze and resolve bottlenecks in inference serving.
- We support deployment and serving of recommendation, ranking, and reinforcement learning models developed by research scientists.
- We apply lightweight model adaptation methods such as LoRA, QLoRA, and PEFT when they fit domain needs.
- We build and maintain deployment pipelines, observability tooling, and tracing infrastructure for agents and serving endpoints.
- We monitor quality regressions, performance issues, and model drift.
- We maintain version control for models, prompts, datasets, and agent configurations.
- We contribute to automated validation, testing, and CI/CD workflows for AI systems.
- We partner with research scientists, backend engineers, and data scientists to bring AI systems into production products.
- We document systems, best practices, and internal tooling.
- We contribute to engineering standards and operational excellence across AI initiatives.
Company:
We are an online gaming company using AI to personalize and enhance player experiences. Our AI Engineering team focuses on bringing AI capabilities into production, with a primary emphasis on agent systems and a secondary focus on model deployment and inference engineering. We do not train foundation models from scratch; instead, we concentrate on production AI systems, model adaptation, inference optimization, and agentic applications. This is an in-person role based in Irvine, CA, and we offer a competitive benefits package that includes health, dental, vision, life insurance, paid time off, parental leave, retirement benefits, and 401(k) matching. The role requires Mandarin Chinese and onsite work, with relocation to Irvine before starting.
Keywords
OrchestrationCUDADeepSpeedOCamlRustPyTorchApache HadoopApache SparkHadoopCI/CDASP.NET
Interested in this role?