← all roles

Applied AI Engineer & Researcher

Interfazeinference provider

San Francisco, CA · senior

Apply at Interfazeoriginal posting on interfaze.ai

posted 7d ago · first seen 7d ago · verified today

Hardwaregpu-generic

Infra layerinference-enginestraining-frameworksml-platform

Role typeresearch-engineersoftware-engineerperformance-engineer

Skills & stackquantizationtensorrt-stackpython-langvllm-enginecudacpp-langrust-lang

Workload stagefine-tuninginferenceevaluationpost-trainingdistributed-inferencetraining-data-infra

All roles

Applied AI Engineer & Researcher

copy markdown

Location: San Francisco, CA

Apply here: https://forms.gle/cSP7f2iy65efyKiq7

About Interfaze

At Interfaze, we're building a new model for deterministic developer tasks that expects high accuracy like OCR, web scraping, classification, STT, and more. We currently power thousands of systems in industries like healthcare, finance, government, SaaS and more which has processed billions of tokens every month.

About this role

We're hiring an AI Engineer & Researcher who is looking to push the boundaries of our models by implementing and experimenting with new research and techniques. As part of the Interfaze lab, you will help improve the performance and capabilities of our models by managing our AI lifecycle, including fine-tuning, automating data collection and cleaning, benchmarking, and deployment.

As a founding member at Interfaze, your role will be dynamic, offering opportunities to spearhead and launch innovative research, papers, and products into production while being the go-to expert for all things AI!

Our tech stack

AreaWhat we use
LanguagesPython, Cython, Typescript, C++ (nice to have), Rust (nice to have)
FrameworksDocker, PyTorch, Transformers, CUDA
Inference enginesSGLang, vLLM, TensorRT-LLM, Ollama, Llama.cpp
Model architecturesGemma 4 Unified, Llama 3.1, Qwen3-VL, GLM 5.2
FrontendNextJS/React
BackendNodeJS with Typescript
Functions & sandboxesModal, Vercel
Model deploymentsGCP, AWS, Azure, Modal
DatabasePostgres (Supabase)
QuerySQL/GraphQL
Payment/BillingStripe

What you will do

  • Research and understand different machine learning techniques and papers, then experiment with real-world implementation
  • Experience with deep learning frameworks (e.g., TensorRT-LLM, PyTorch), and AI development tools is a must
  • Working closely with the team to train, deploy and serve models at scale on popular cloud providers like AWS, GCP, and more
  • Write papers on your research and experiments contributing to the growing open-source AI world
  • Writing detailed benchmarks on comparisons and analysis (e.g. comparing performance for two embedding models or benchmarking whisper 3 with another model)
  • Quantization and fine-tuning of LLMs and other models for our developer use cases
  • Optimizing models and GPU infra for cost-to-scale ratio based on load
  • Writing Python code and Jupyter Notebooks

Who you are

  • A strong AI Researcher with at least 5 years of experience, and paper publications at venues like ICML, NeurIPS, ACL, NAACL, and IEEE
  • Worked with most of our GPU infra and have a good understanding of our tech stack
  • Comfortable with working across the stack - model deployment, benchmarking, fine-tuning, monitoring, and scaling the service
  • Up-to-date with the latest papers, techniques, and experiments
  • Using AI to 100x your productivity
  • Natural curiosity to experiment with new concepts and tools while working closely with our community of developers
  • You care about building great products and love to geek out about the latest tech in AI
  • Open to feedback and have a growth mindset

Bonus if you

  • Enjoy working in a startup environment where you can take initiative and contribute to the zero-to-one phase of development
  • Previously ran a startup or worked as a founding team member in a startup

Benefits

  • Great compensation package and equity/shares
  • Flexible and remote-friendly
  • Annually run off-sites
  • Learn and Grow - we provide mentorship and send you to events that help you build your network and skills
  • Annual Education Allowance
  • Let's build benefits together

Process

  • Apply here: https://forms.gle/cSP7f2iy65efyKiq7
  • The entire process is fully remote and all communication will happen over email or via video chat.
  • Once you've submitted your application, the team will review your submission and may reach out for a short screening interview over a video call.
  • If you pass the screen you may be invited to complete a short task or discuss your past work.
  • If you pass the task/screen you will be invited to up to 2 follow-up interviews. The calls:
    • usually take between 20-45 minutes each depending on the interviewer.
    • are all 1:1.
    • will be with both founders, a member of either the growth or engineering team (depending on the role) and usually one other person from your immediate team or function.

Apply for this role ->

All roles