Staff Software Engineer - GenAI inference
Databricks · San Francisco, California
About this role
P-1285
About This Role
As a staff software engineer for GenAI inference, you will lead the architecture, development, and optimization of the inference engine that powers Databricks Foundation Model API.. You’ll bridge research advances and production demands, ensuring high throughput, low latency, and robust scaling. Your work will encompass the full GenAI inference stack: kernels, runtimes, orchestration, memory, and integration with frameworks and orchestration systems.
What You Will Do
Own and drive the architecture, design, and implementation of the inference engine, and collaborate on model-serving stack optimized for large-scale LLMs inference
Partner closely with researchers to bring new model architectures or features (sparsity, activation compression, mixture-of-experts) into the engine
Lead the end-to-end optimization for latency, throughput, memory efficiency, and hardware utilization across GPUs, and accelerators
Define and guide standards to build and maintain instrumentation, profiling, and tracing tooling to uncover bottlenecks and guide optimizations…
Summary from Databricks's official Greenhouse career feed — read the full description on the original posting ↗
The full job description lives on Databricks's official careers page. AI Stack Jobs links you straight to it — we never sit between you and the employer.
More AI/ML roles at Databricks
- 2026-08-17 Remote - California; Remote - Colorado; Remote - Oregon; Remote - Washington H-1B SponsorRemote
- 2026-08-17 Bengaluru, India H-1B Sponsor
- 2026-08-14 United States H-1B Sponsor
- 2026-08-14 Seattle, Washington H-1B Sponsor
- 2026-08-14 San Francisco, California H-1B Sponsor
- 2026-08-14 Bengaluru, India H-1B Sponsor