HHiring Reality
← KRAFTON

[AI Research Div.] ML Serving Engineer (5년 이상)

KRAFTON
Location
Seoul
Posted
1 month ago
Department
Data
What they actually want (must-haves)
  • 5+ years of experience in designing, building, and operating AI/LLM model serving systems in production environments
  • Experience applying high-performance LLM inference engines like vLLM or TensorRT-LLM for system optimization
  • Knowledge of memory and computation efficiency techniques such as KV Cache Offloading/Loading and quantization
  • Strong analytical skills to address latency and performance issues across inference engines and GPU hardware
  • Experience building and enhancing monitoring and observability systems for serving stability and CI/CD automation
  • Ability to travel internationally without restrictions
Nice to have
  • Experience in making architectural decisions and understanding trade-offs in performance, cost, and stability
  • Contributions to open-source serving frameworks like vLLM or LMCache
  • Deep understanding of high-performance LLM inference stacks and experience in custom kernel development and CUDA acceleration
What the job really is

The ML Serving Engineer will design and implement high-performance ML serving architectures for game applications, ensuring stable and efficient operation of AI models in production environments. Responsibilities include building real-time serving infrastructure for large language models (LLMs), optimizing performance, and establishing CI/CD pipelines for continuous deployment and monitoring.

Things to weigh
  • No salary or specific benefits mentioned
  • The role requires a significant level of expertise and experience, which may limit applicants
  • The position involves a probation period of 5 months with no changes in employment type or salary during this time
Job score3/5
Benefits1/5
Freshness4/5
Career value5/5
Role clarity5/5
Pay transparency0/5

Applying to KRAFTON?

See how your résumé matches this role — and tailor it from what actually gets interviews.