Skip to main content
Craftware
Scraped fromWorkYesterday
Data Science

Data Scientist – RAG & Document Intelligence

PythonRagLlamaindexLangchainLightragRagasDeepevalLangfuseLanggraphPydantic AiPgvectorPineconeQdrantWeaviateAzureAWSDatabricksDelta LakeUnity CatalogMlflowFastAPIModel Context Protocol
Work Type
Remote
Job Type
Full Time
Location
Worldwide
Salary
PLN 160–190 / HOUR

About the Position

Craftware is a technology company of over 500 experts, empowering large organizations to solve complex business challenges with modern IT solutions. They are building a platform that lets business users retrieve, synthesize, and act on information locked inside complex enterprise documents using plain language.

Responsibilities

  • Design, experiment with, and continuously optimize RAG pipelines – chunking strategies, embedding models, hybrid search, re-ranking, and context assembly
  • Benchmark and evaluate LLMs, embedding models, and document parsing solutions across accuracy, latency, reliability, and cost
  • Build evaluation datasets and define RAG-specific quality metrics using frameworks like RAGAS, DeepEval, and Langfuse
  • Design and iterate on multi-agent system architectures (LangGraph, LangChain, Pydantic AI)
  • Handle diverse, noisy document formats – PDFs, Word docs, presentations, scanned files, tables, charts, and mixed-format corpora
  • Track token consumption, latency profiles, and retrieval quality – making principled trade-off decisions between capability and operational cost
  • Shape a system that will serve business users across commercial, marketing, R&D, and product supply globally

Requirements

  • Strong Python skills with hands-on experience in ML and generative AI workflows, including production-grade code
  • Solid NLP background – text representation, semantic search, embeddings, language model behaviour
  • Deep hands-on experience designing and optimizing RAG pipelines – chunking, hybrid search, re-ranking (LlamaIndex, LangChain, LightRAG)
  • Experience with document parsing across diverse and noisy formats, including tables, charts, and figures
  • Experience with LLM evaluation frameworks and RAG-specific quality metrics (RAGAS, DeepEval, Langfuse)
  • Familiarity with multi-agent AI frameworks – LangGraph, LangChain, or Pydantic AI
  • Experience with vector databases (pgvector, Pinecone, Qdrant, Weaviate)
  • Experience with Azure cloud services (Apps, Containers, Storage, AI Search, AI Foundry) and/or AWS equivalents
  • Experience with Databricks (Delta Lake, Unity Catalog, MLFlow)
  • Structured, experiment-driven approach to problem solving
  • Fluent English – written and spoken

Benefits

  • B2B contract
  • Daily support from team leaders
  • Dedicated certification budget
  • Assistance in defining and support in your development path
  • Benefits package
  • Integration trips/events
Data Scientist – RAG & Document IntelligencePLN 160–190 / HOUR
View Original