Skip to main content
Neurons Lab
Scraped fromIndeedYesterday
Frontend

Data Engineer | Data Scientist — Unstructured Data & AI Pipelines

PythonSQLApache AirflowStep FunctionsAWSGCPpgvectorOpenSearchPineconeGoogle WorkspaceMicrosoft 365SlackCRMMCPComposioLLMRAG
Work Type
-
Job Type
Part Time
Location
Worldwide
Salary
Not specified

About the Position

Join Neurons Lab as a Data Engineer on a part-time engagement with a European private investment group. The project involves building unstructured data pipelines for ingestion, entity resolution, embedding, and retrieval infrastructure, using Python, Airflow, and cloud services.

Responsibilities

  • Stand up capture by default: notetaker on every call with speaker attribution, plus ingestion from mail, Slack and messengers — designed as opt-out, not opt-in, and reversible if the client changes their mind.
  • Backfill the archive: years of historical email, Slack, board protocols, decks and portfolio updates — parsed, deduplicated and dated correctly.
  • Build document parsing for the awkward long tail: PDFs, scanned board packs, spreadsheets, slide decks, forwarded attachments.
  • Implement identity / entity resolution: the same person across Slack handle, mail alias and calendar invite; the same portfolio company across a deck, a mail thread and a CRM record.
  • Build chunking and embedding pipelines and load the vector + graph stores behind the ontology the architect defines.
  • Implement incremental sync through the connector layer (MCP / Composio-class) — no full re-crawls, no silent drift, clear handling of edits and deletions.
  • Attach access scope and provenance to every record at ingestion, so permission-aware retrieval and audit are possible downstream rather than bolted on.
  • Run PII detection, redaction and retention logic; evidence to the client's security function what is stored, where, and for how long.
  • Orchestrate with Airflow / Step Functions; build repeatable, monitored pipelines rather than scripts, with alerting when a source stops flowing.
  • Keep cost and latency under control at volume — batching, incremental embedding, storage tiering — and report the unit economics.
  • Write runbooks so the client's own team can operate this after handover.

Requirements

  • Strong Python and solid SQL
  • Unstructured-data pipelines: transcripts, mail, chat, documents — parsing, normalisation, deduplication
  • Embedding / retrieval infrastructure: chunking strategies, vector stores (pgvector, OpenSearch, Pinecone-class), plus loading a graph store
  • API and connector integration at scale: Google Workspace / M365, Slack, CRM; rate limits, pagination, incremental cursors, webhooks
  • Entity resolution / record linkage (deterministic + fuzzy) without a clean shared key
  • Orchestration: Airflow, Step Functions or equivalent; idempotent, restartable jobs
  • AWS and/or GCP data stack; comfortable in a private / VPC deployment
  • PII detection, redaction, encryption and retention in practice
  • Clear written English; documents for handover and works well async in a small distributed pod
  • 4+ years in data engineering, with real unstructured / semi-structured work (not only warehouse modelling)
  • Demonstrated experience integrating many third-party APIs into one coherent store, including historical backfill
  • Experience building pipelines feeding an LLM / retrieval system — strong plus
  • Experience handling sensitive personal data in a regulated or security-sensitive environment
  • Comfortable being the only data engineer on a small (2.5-FTE) pod, at part-time allocation, without hand-holding

Who to contact

Data Engineer | Data Scientist — Unstructured Data & AI Pipelines
View Original