N
Neurons Lab
Frontend
Data Engineer | Data Scientist — Unstructured Data & AI Pipelines
PythonSQLApache AirflowStep FunctionsAWSGCPpgvectorOpenSearchPineconeGoogle WorkspaceMicrosoft 365SlackCRMMCPComposioLLMRAG
Про позицію
Join Neurons Lab as a Data Engineer on a part-time engagement with a European private investment group. The project involves building unstructured data pipelines for ingestion, entity resolution, embedding, and retrieval infrastructure, using Python, Airflow, and cloud services.
Обовʼязки
- Stand up capture by default: notetaker on every call with speaker attribution, plus ingestion from mail, Slack and messengers — designed as opt-out, not opt-in, and reversible if the client changes their mind.
- Backfill the archive: years of historical email, Slack, board protocols, decks and portfolio updates — parsed, deduplicated and dated correctly.
- Build document parsing for the awkward long tail: PDFs, scanned board packs, spreadsheets, slide decks, forwarded attachments.
- Implement identity / entity resolution: the same person across Slack handle, mail alias and calendar invite; the same portfolio company across a deck, a mail thread and a CRM record.
- Build chunking and embedding pipelines and load the vector + graph stores behind the ontology the architect defines.
- Implement incremental sync through the connector layer (MCP / Composio-class) — no full re-crawls, no silent drift, clear handling of edits and deletions.
- Attach access scope and provenance to every record at ingestion, so permission-aware retrieval and audit are possible downstream rather than bolted on.
- Run PII detection, redaction and retention logic; evidence to the client's security function what is stored, where, and for how long.
- Orchestrate with Airflow / Step Functions; build repeatable, monitored pipelines rather than scripts, with alerting when a source stops flowing.
- Keep cost and latency under control at volume — batching, incremental embedding, storage tiering — and report the unit economics.
- Write runbooks so the client's own team can operate this after handover.
Вимоги
- Strong Python and solid SQL
- Unstructured-data pipelines: transcripts, mail, chat, documents — parsing, normalisation, deduplication
- Embedding / retrieval infrastructure: chunking strategies, vector stores (pgvector, OpenSearch, Pinecone-class), plus loading a graph store
- API and connector integration at scale: Google Workspace / M365, Slack, CRM; rate limits, pagination, incremental cursors, webhooks
- Entity resolution / record linkage (deterministic + fuzzy) without a clean shared key
- Orchestration: Airflow, Step Functions or equivalent; idempotent, restartable jobs
- AWS and/or GCP data stack; comfortable in a private / VPC deployment
- PII detection, redaction, encryption and retention in practice
- Clear written English; documents for handover and works well async in a small distributed pod
- 4+ years in data engineering, with real unstructured / semi-structured work (not only warehouse modelling)
- Demonstrated experience integrating many third-party APIs into one coherent store, including historical backfill
- Experience building pipelines feeding an LLM / retrieval system — strong plus
- Experience handling sensitive personal data in a regulated or security-sensitive environment
- Comfortable being the only data engineer on a small (2.5-FTE) pod, at part-time allocation, without hand-holding
До кого писати
Data Engineer | Data Scientist — Unstructured Data & AI Pipelines
Оригінал