Senior AI Engineer
Details
Über die Stelle
We are growing the development team behind one of our client's most successful products — an integrated delivery platform built on the depth of real-world engagement experience. It enables cross-team collaboration, real-time transparency, and better decision-making.
Your job is the AI layer of that product: end-to-end LLM pipelines, retrieval, and multi-agent orchestration running in production, not in a demo. Expect deep Python work, real evaluation metrics, and full ownership of how the system behaves under load, cost, and latency constraints.
Ihre Aufgaben
- LLM pipeline architecture: design end-to-end pipelines using RAG, embeddings, and orchestration layers
- Multi-agent systems: build modular, scalable agent workflows in LangChain, LangGraph, CrewAI, or AutoGen, including observer and fallback agents
- Retrieval engineering: build chunking and ingestion pipelines for PDFs and unstructured data, and deploy vector stores with semantic search and reranking
- Model strategy: evaluate model options against performance, cost, and capability trade-offs, and define embedding strategies, context windows, and prompt structures
- Evaluation & QA: define and track metrics for GenAI and RAG systems — faithfulness, precision, recall, semantic similarity — using RAGAS, sklearn, and NumPy
- Observability: instrument the system with LangFuse, Datadog, or custom tracing to monitor, debug, and optimize
- Delivery engineering: maintain CI/CD pipelines, Docker containers, test coverage, and API integrations
- Communication: turn complex models into clear insights for executive stakeholders, lead design sessions, and mentor other developers
Erforderliche Fähigkeiten
- 8+ years in software engineering, with at least 3 years focused on AI/ML, NLP, or LLM-based applications
- Expert-level Python — modular, class-based, production-grade code you can debug across distributed services
- Proven GenAI delivery: prompt engineering, embedding-based retrieval, and multi-agent orchestration patterns in production
- Modern AI stack: LangChain, LangGraph, CrewAI, AutoGen, Hugging Face
- Vector databases: Azure AI Search, FAISS, or Postgres with pgvector
- Engineering foundations: Jupyter, Docker, GitHub, REST API integration, CI/CD via GitHub Actions
- AI-assisted development: comfortable with Cursor, GitHub Copilot, and similar tools without sacrificing security or maintainability
- Preferred: multimodal LLMs, ReAct-style prompting, planner-executor agents, human-in-the-loop RAG evaluation
Was erwartet Sie?
- A production product with real users — no proof-of-concept purgatory
- Full ownership of the AI architecture, from retrieval strategy to observability
- A 12-month engagement with a strong likelihood of extension
- Fully remote work with a team that meets in person roughly once a quarter in Prague
- Direct exposure to executive stakeholders and the reasoning behind product decisions
Startdatum
- Start: ASAP
- Workload: Full-time, 12-month contract with a strong likelihood of extension
- Setup: fully remote, occasional in-person team sessions in Prague (approx. 1× per quarter)
- Overlap: minimum 14:00–18:00 CET with the US team
- Selection process: includes a HackerRank challenge
Prämie
Lassen Sie sich belohnen.
Keine langen Formulare – nur Name und Kontakt. Wir kümmern uns um den Rest. Die Prämie wird nach bestandener dreimonatiger Probezeit ausgezahlt.

Prozess
Vier Schritte, keine Hausaufgaben.
5 Minuten
Bewerben
Lebenslauf oder LinkedIn, kein Anschreiben. Wir antworten innerhalb von 48 Stunden – jedem, auch bei einer Absage.
FRAGEN
— welcher Stack
— wie viele Personen
— seit wann
— was Sie lösen
45 Minuten
Technisches Gespräch
Mit einem Ingenieur, nicht mit einem Recruiter. Architektur, Kompromisse, Ihre echten Projekte. Kein „Beschreiben Sie eine Situation, in der Sie...“.
FRAGEN
— welcher Stack
— wie viele Personen
— seit wann
— was Sie lösen
90 Minuten
Pairing an echtem Code
Ein bestehendes Repo, ein echter Bug oder ein kleines Feature. Uns interessiert, wie Sie denken und debuggen – nicht Whiteboard-Algorithmen.
FRAGEN
— welcher Stack
— wie viele Personen
— seit wann
— was Sie lösen
Innerhalb von 7 Tagen
Angebot
Eine konkrete Zahl, ein konkretes Projekt, ein konkretes Team. Entscheidung innerhalb einer Woche nach der Pair-Session.
FRAGEN
— welcher Stack
— wie viele Personen
— seit wann
— was Sie lösen




