Skip to content
NewFull-timeRemotePosted today

Senior ML Engineer (Token Factory)

Nebius · Remote

#ml

About the role

<div class=&quot;content-intro&quot;><p><strong>About Nebius:</strong></p> <p>Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.</p> <p>Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.</p> <p>Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&amp;D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&amp;D.</p></div><h2 id=&quot;The-role&quot; data-renderer-start-pos=&quot;1&quot;><strong data-renderer-mark=&quot;true&quot;>The role</strong></h2> <p data-renderer-start-pos=&quot;11&quot;>Token Factory is a part of Nebius Cloud, one of the world’s largest GPU clouds, running tens of thousands of GPUs. We are building an inference &amp; fine-tuning platform that makes every kind of foundation model — text, vision, audio, and emerging multimodal architectures — fast, reliable, and effortless to train &amp; deploy at massive scale.</p> <div class=&quot;ewa-rteLine&quot;><strong>Some directions we currently working on and which you can be a part of:</strong></div> <ul class=&quot;ak-ul&quot; data-indent-level=&quot;1&quot;> <li> <div class=&quot;ewa-rteLine&quot;> <p><strong>Advanced Fine-Tuning:</strong> Enhancing fine-tuning methodologies - both LoRA-based and full-parameter - for cutting-edge LLMs (e.g., GPT-OSS, Kimi K2.5, DeepSeek V3.1/V3.2, GLM-4.7), focusing on both model quality and training efficiency.</p> </div> </li> <li> <div class=&quot;ewa-rteLine&quot;><strong>Inference Optimization:</strong> Identifying LLM inference bottlenecks to drive production speedups. This involves building model training and evaluation pipelines in JAX for speculative decoding, experimenting with architectures (dense/MoE, auto-regressive/parallel), and deriving scaling laws to guide resource allocation.&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;</div> </li> <li><strong>Low&nbsp;</strong><strong>Precision Training &amp; Inference:</strong> Investigating low-precision (FP8, NVFP4/MXFP4) methodologies for supervised fine-tuning and reinforcement learning - spanning both inference and training - optimized for modern hardware</li> </ul> <p><strong data-renderer-mark=&quot;true&quot;>We expect you to have:</strong></p> <ul class=&quot;ak-ul&quot; data-indent-level=&quot;1&quot;> <li> <p data-renderer-start-pos=&quot;927&quot;>A profound understanding of theoretical foundations of machine learning and reinforcement learning.</p> </li> <li> <p data-renderer-start-pos=&quot;1030&quot;>Deep expertise in modern deep learning for language processing and generation</p> </li> <li> <p data-renderer-start-pos=&quot;1111&quot;>Experience with training large models on multiple computational nodes</p> </li> <li> <p data-renderer-start-pos=&quot;1196&quot;>Reasonable understanding of performance aspects of large neural network training (sharding strategies, custom kernels, hardware features etc.)</p> </li> <li> <p data-renderer-start-pos=&quot;1342&quot;>Strong software engineering skills (we mostly use Python)</p> </li> <li> <p data-renderer-start-pos=&quot;1403&quot;>Deep experience with modern deep learning frameworks (we use JAX)</p> </li> <li> <p data-renderer-start-pos=&quot;1472&quot;>Proficiency in contemporary software engineering approaches, including CI/CD, version control and unit testing</p> </li> <li> <p data-renderer-start-pos=&quot;1586&quot;>Strong communication and leadership abilities</p> </li> </ul> <p><strong data-renderer-mark=&quot;true&quot;>Nice to have:</strong></p> <ul class=&quot;ak-ul&quot; data-indent-level=&quot;1&quot;> <li> <p data-renderer-start-pos=&quot;1652&quot;><strong data-renderer-mark=&quot;true&quot;>Previous experience working with language models or other similar NLP technologies.</strong></p> </li> <li> <p data-renderer-start-pos=&quot;1739&quot;>Familiarity with important ideas in LLM space, such as MHA, RoPE, ZeRO/FSDP, Flash Attention, quantization</p> </li> <li> <p data-renderer-start-pos=&quot;1849&quot;>A track record of building and delivering products (not necessarily ML-related) in a dynamic startup-like environment.</p> </li> <li> <p data-renderer-start-pos=&quot;1971&quot;>Strong engineering skills, including experience in developing large distributed systems or high-load web services.</p> </li> <li> <p data-renderer-start-pos=&quot;2089&quot;>Open-source projects that showcase your engineering prowess</p> </li> <li> <p data-renderer-start-pos=&quot;2152&quot;>Excellent command of the English language, alongside superior writing, articulation, and communication skills.</p> </li> </ul><div class=&quot;content-conclusion&quot;><p><strong>Benefits &amp; Perks:</strong></p> <ul> <li>Competitive compensation</li> <li>Career growth and learning opportunities</li> <li>Flexibility and ownership</li> <li>Collaborative and innovative culture</li> <li>Opportunity to work on impactful AI projects</li> <li>International environment and talented teams</li> </ul> <p><strong>What&#39;s it like to work at Nebius:</strong></p> <p>Fast moving&nbsp;- Bold thinking&nbsp;- Constant growth&nbsp;- Meaningful impact&nbsp;- Trust and real ownership&nbsp;- Opportunity to shape the future of AI&nbsp;</p> <p><strong>Equal Opportunity Statement:</strong></p> <p>Nebius is an equal opportunity employer. We are committed to fostering an incl…

Apply for this role

Apply right here on 9180 — we'll forward your application to Nebius. No external redirects, no lost candidates.

1 · About you
Your work

2 · Your work

🔒 Never shown to anyone — used only in anonymous, aggregate Bangalore pay benchmarks.

One link is enough — your best work, or a résumé the team can open.

Your pitch

3 · Your pitch

Free · forwarded straight to the hiring team.

Originally listed on Arbeitnow. View the original posting ↗

More roles like this

All jobs →