Staff Software Engineer - GenAI inference

Databricks · San Francisco, California

Company
Databricks
Location
San Francisco, California
Last updated
2026-08-11
Source
Official Greenhouse feed

About this role

P-1285

About This Role

As a staff software engineer for GenAI inference, you will lead the architecture, development, and optimization of the inference engine that powers Databricks Foundation Model API.. You’ll bridge research advances and production demands, ensuring high throughput, low latency, and robust scaling. Your work will encompass the full GenAI inference stack: kernels, runtimes, orchestration, memory, and integration with frameworks and orchestration systems.

What You Will Do

Own and drive the architecture, design, and implementation of the inference engine, and collaborate on model-serving stack optimized for large-scale LLMs inference

Partner closely with researchers to bring new model architectures or features (sparsity, activation compression, mixture-of-experts) into the engine

Lead the end-to-end optimization for latency, throughput, memory efficiency, and hardware utilization across GPUs, and accelerators

Define and guide standards to build and maintain instrumentation, profiling, and tracing tooling to uncover bottlenecks and guide optimizations…

Summary from Databricks's official Greenhouse career feed — read the full description on the original posting

Apply on Databricks's site

The full job description lives on Databricks's official careers page. AI Stack Jobs links you straight to it — we never sit between you and the employer.

More AI/ML roles at Databricks

  1. 2026-08-17 Remote - California; Remote - Colorado; Remote - Oregon; Remote - Washington H-1B SponsorRemote
  2. 2026-08-17 Bengaluru, India H-1B Sponsor
  3. 2026-08-14 United States H-1B Sponsor
  4. 2026-08-14 Seattle, Washington H-1B Sponsor
  5. 2026-08-14 San Francisco, California H-1B Sponsor
  6. 2026-08-14 Bengaluru, India H-1B Sponsor