Senior AI Engineer
Details
About job
We are growing the development team behind one of our client's most successful products — an integrated delivery platform built on the depth of real-world engagement experience. It enables cross-team collaboration, real-time transparency, and better decision-making.
Your job is the AI layer of that product: end-to-end LLM pipelines, retrieval, and multi-agent orchestration running in production, not in a demo. Expect deep Python work, real evaluation metrics, and full ownership of how the system behaves under load, cost, and latency constraints.
Your Responsibilities
- LLM pipeline architecture: design end-to-end pipelines using RAG, embeddings, and orchestration layers
- Multi-agent systems: build modular, scalable agent workflows in LangChain, LangGraph, CrewAI, or AutoGen, including observer and fallback agents
- Retrieval engineering: build chunking and ingestion pipelines for PDFs and unstructured data, and deploy vector stores with semantic search and reranking
- Model strategy: evaluate model options against performance, cost, and capability trade-offs, and define embedding strategies, context windows, and prompt structures
- Evaluation & QA: define and track metrics for GenAI and RAG systems — faithfulness, precision, recall, semantic similarity — using RAGAS, sklearn, and NumPy
- Observability: instrument the system with LangFuse, Datadog, or custom tracing to monitor, debug, and optimize
- Delivery engineering: maintain CI/CD pipelines, Docker containers, test coverage, and API integrations
- Communication: turn complex models into clear insights for executive stakeholders, lead design sessions, and mentor other developers
Skills required
- 8+ years in software engineering, with at least 3 years focused on AI/ML, NLP, or LLM-based applications
- Expert-level Python — modular, class-based, production-grade code you can debug across distributed services
- Proven GenAI delivery: prompt engineering, embedding-based retrieval, and multi-agent orchestration patterns in production
- Modern AI stack: LangChain, LangGraph, CrewAI, AutoGen, Hugging Face
- Vector databases: Azure AI Search, FAISS, or Postgres with pgvector
- Engineering foundations: Jupyter, Docker, GitHub, REST API integration, CI/CD via GitHub Actions
- AI-assisted development: comfortable with Cursor, GitHub Copilot, and similar tools without sacrificing security or maintainability
- Preferred: multimodal LLMs, ReAct-style prompting, planner-executor agents, human-in-the-loop RAG evaluation
What can you expect?
- A production product with real users — no proof-of-concept purgatory
- Full ownership of the AI architecture, from retrieval strategy to observability
- A 12-month engagement with a strong likelihood of extension
- Fully remote work with a team that meets in person roughly once a quarter in Prague
- Direct exposure to executive stakeholders and the reasoning behind product decisions
Start date
- Start: ASAP
- Workload: Full-time, 12-month contract with a strong likelihood of extension
- Setup: fully remote, occasional in-person team sessions in Prague (approx. 1× per quarter)
- Overlap: minimum 14:00–18:00 CET with the US team
- Selection process: includes a HackerRank challenge
Reward
Get rewarded.
No lengthy forms — just a name and a contact. We handle the rest. Reward paid once they pass their three-month probation.

Process
Four steps, no take-home assignments.
5 minutes
Apply
CV or LinkedIn, no cover letter. We reply within 48 hours — to everyone, including the no's.
ASK
— which stack
— how many people
— from when
— what you're solving
45 minutes
Tech call
With an engineer, not a recruiter. Architecture, tradeoffs, your real projects. No "describe a situation where you had to…".
ASK
— which stack
— how many people
— from when
— what you're solving
90 minutes
Pair na reálném kódu
An existing repo, a real bug or a small feature. We care how you think and debug — not whiteboard algorithms.
ASK
— which stack
— how many people
— from when
— what you're solving
Within 7 days
Offer
A concrete number, a concrete project, a concrete team. Decision within a week of the pair session.
ASK
— which stack
— how many people
— from when
— what you're solving



