AI Software Engineer
The job description
Tech stack. Python, LLM APIs (OpenAI, Anthropic), RAG architectures, vector databases, prompt engineering, agent frameworks, evaluation harnesses, TypeScript, guardrails
About the role
You will build AI-powered product features using large language models and related technologies, working at the frontier where product intuition matters as much as engineering skill. AI software engineers here design the complete systems around the models: retrieval pipelines, prompt architectures, evaluation harnesses, safety guardrails, and the user experiences that make AI genuinely useful rather than merely impressive in a demo. You will iterate fast on prompts and architectures while maintaining the engineering discipline that keeps AI systems reliable, safe, and cost-effective in production. Shipping AI that users trust daily is the standard.
What you will achieve
- Ship AI product features to production: chat experiences, copilots, or automation tools that users genuinely rely on in their daily workflows
- Build RAG systems with high answer quality: document ingestion pipelines, chunking strategy, retrieval tuning, and citation-backed responses users can verify
- Deliver rigorous AI evaluation: golden datasets, automated quality scoring, human review loops, and regression detection across model version updates
- Control AI costs and latency through smart model selection, caching strategies, prompt optimization, and efficient system architectures
- Implement safety guardrails: content filtering, hallucination mitigation techniques, and graceful failure modes that protect users and the business
What you will bring
Must-haves
- 2 to 5 years of software engineering with hands-on experience building LLM-powered applications that reached production users
- Deep familiarity with LLM APIs from OpenAI, Anthropic, or open models: capabilities, limitations, pricing structures, and failure modes
- Experience with RAG architectures: embeddings selection, vector databases such as Pinecone, pgvector, or Qdrant, and retrieval quality tuning
- Understanding of prompt engineering as an engineering discipline: system prompts, few-shot examples, structured outputs, and version control for prompts
- Knowledge of AI evaluation methodology: building representative test sets, defining quality metrics, and measuring improvements rigorously
- Awareness of AI safety and reliability concerns: hallucinations, prompt injection attacks, data privacy, and appropriate human oversight design
- BS in Computer Science or equivalent experience
Nice-to-haves
- Experience fine-tuning open models with LoRA or full fine-tuning for specific tasks or specialized domains
- Familiarity with agent frameworks: tool use patterns, multi-step reasoning orchestration, and error recovery strategies
- Knowledge of multimodal AI: vision, speech, or document understanding integrated into cohesive products
- Experience with AI observability: tracing prompts, monitoring output quality, and debugging AI behavior in production
Google
Meta
Apple
Microsoft
Amazon
Oracle
Netflix
NVIDIA