Senior AI Engineer

Can you take LLM pipelines from notebook to production?
AI / ML
Remote
Full-time
//

Podrobnosti

O pozici

We are growing the development team behind one of our client's most successful products — an integrated delivery platform built on the depth of real-world engagement experience. It enables cross-team collaboration, real-time transparency, and better decision-making.

Your job is the AI layer of that product: end-to-end LLM pipelines, retrieval, and multi-agent orchestration running in production, not in a demo. Expect deep Python work, real evaluation metrics, and full ownership of how the system behaves under load, cost, and latency constraints.

Vaše odpovědnost

  • LLM pipeline architecture: design end-to-end pipelines using RAG, embeddings, and orchestration layers
  • Multi-agent systems: build modular, scalable agent workflows in LangChain, LangGraph, CrewAI, or AutoGen, including observer and fallback agents
  • Retrieval engineering: build chunking and ingestion pipelines for PDFs and unstructured data, and deploy vector stores with semantic search and reranking
  • Model strategy: evaluate model options against performance, cost, and capability trade-offs, and define embedding strategies, context windows, and prompt structures
  • Evaluation & QA: define and track metrics for GenAI and RAG systems — faithfulness, precision, recall, semantic similarity — using RAGAS, sklearn, and NumPy
  • Observability: instrument the system with LangFuse, Datadog, or custom tracing to monitor, debug, and optimize
  • Delivery engineering: maintain CI/CD pipelines, Docker containers, test coverage, and API integrations
  • Communication: turn complex models into clear insights for executive stakeholders, lead design sessions, and mentor other developers

Požadované dovednosti

  • 8+ years in software engineering, with at least 3 years focused on AI/ML, NLP, or LLM-based applications
  • Expert-level Python — modular, class-based, production-grade code you can debug across distributed services
  • Proven GenAI delivery: prompt engineering, embedding-based retrieval, and multi-agent orchestration patterns in production
  • Modern AI stack: LangChain, LangGraph, CrewAI, AutoGen, Hugging Face
  • Vector databases: Azure AI Search, FAISS, or Postgres with pgvector
  • Engineering foundations: Jupyter, Docker, GitHub, REST API integration, CI/CD via GitHub Actions
  • AI-assisted development: comfortable with Cursor, GitHub Copilot, and similar tools without sacrificing security or maintainability
  • Preferred: multimodal LLMs, ReAct-style prompting, planner-executor agents, human-in-the-loop RAG evaluation

Co můžete očekávat?

  • A production product with real users — no proof-of-concept purgatory
  • Full ownership of the AI architecture, from retrieval strategy to observability
  • A 12-month engagement with a strong likelihood of extension
  • Fully remote work with a team that meets in person roughly once a quarter in Prague
  • Direct exposure to executive stakeholders and the reasoning behind product decisions

Datum zahájení

  • Start: ASAP
  • Workload: Full-time, 12-month contract with a strong likelihood of extension
  • Setup: fully remote, occasional in-person team sessions in Prague (approx. 1× per quarter)
  • Overlap: minimum 14:00–18:00 CET with the US team
  • Selection process: includes a HackerRank challenge

contact_us.ts
Kontaktujte nás
Maximální velikost souboru 10 MB.
Nahrávání...
fileuploaded.jpg
Upload failed. Max size for files is 10 MB.
// odpovíme do 24 hodin
Děkujeme! Vaše zpráva byla přijata.
Jejda! Při odesílání formuláře došlo k chybě.
//

Odměna

Doporučte IT specialistu.
Získejte odměnu.

Žádné zdlouhavé formuláře – stačí jméno a kontakt. O zbytek se postaráme my. Odměnu vyplatíme po úspěšném absolvování tříměsíční zkušební doby.

Fotografie zakladatele společnosti
Jovana Cvetkova
Recruitment Consultant
//

Proces

Čtyři kroky, žádné úkoly domů.

1
//

5 minut

Odeslat

CV nebo LinkedIn, motivační dopis není třeba. Odpovíme do 48 hodin — všem, i těm, kterým řekneme ne.

//

ZEPTAT SE

— jaký stack

— kolik lidí

— od kdy

— co řešíte

2
//

45 minut

Technický hovor

S inženýrem, ne s personalistou. Architektura, kompromisy, vaše reálné projekty. Žádné „popište situaci, kdy jste museli…“.

//

ZEPTAT SE

— jaký stack

— kolik lidí

— od kdy

— co řešíte

3
//

90 minut

Párové programování na reálném kódu

Existující repozitář, skutečná chyba nebo malá funkce. Záleží nám na tom, jak přemýšlíte a ladíte kód – ne na algoritmických hádankách u tabule.

//

ZEPTAT SE

— jaký stack

— kolik lidí

— od kdy

— co řešíte

4
//

Do 7 dnů

Nabídka

Konkrétní částka, konkrétní projekt, konkrétní tým. Rozhodnutí do týdne od společného setkání.

//

ZEPTAT SE

— jaký stack

— kolik lidí

— od kdy

— co řešíte