Data Platform Engineer – Data Operations (all genders)
Stark · Munich
What this role requires
4 requirements, read out of the advert rather than guessed from the job title:
Also mentioned, not required: Docker, CI/CD, GCS, S3, English. Worth having, but their absence is not what gets a CV filtered out.
See how often each of these is required across open data roles in Europe.
Free, no card. Tells you which of them your CV evidences and which it only implies.
Job description
About Us STARK is a new kind of defence technology company revolutionizing the way autonomous systems are deployed across multiple domains. We design, develop and manufacture high-performance unmanned systems that are software-defined, mass-scalable, and cost-effective. This provides our operators with a decisive edge in highly contested environments.
We're focused on delivering deployable, high-performance systems — not future promises. In a time of rising threats, STARK is bolstering the technological edge of NATO Allies and their Partners to deter aggression and defend Europe — today.
About the team The Data Operations team owns the entire data lifecycle behind STARK's AI stack: collection, acquisition, generation, curation, and management. We run our own data-collection campaigns across Europe, evaluate new sensors and platforms, and build the internal data platform that turns raw recordings into ready-to-use datasets. Everything we produce feeds directly into the perception and autonomy systems deployed on STARK's platforms — a real data advantage is built, not bought. The team is scaling up right now: real scope, direct impact, no legacy.
Your mission Data is the fuel of STARK’s AI stack — you build the engine that makes it usable. You own the software backbone of our data platform: the metadata systems, ETL pipelines, data contracts, catalogs, databases, and internal tools that let engineers find, understand, validate, and reuse terabytes of multi-sensor field data in minutes, not days. You treat data context as a product: structured, searchable, version-aware, documented, and traceable from raw recording to processed asset, annotation delivery, dataset, and downstream ML workflow. Today, much of this is manual, scattered, or implicit — your job is to automate it away, support labeling efforts with the right data tooling, and turn operational data into reliable systems.
Responsibilities Design, implement, and maintain our metadata database and data catalog (datasets, recordings, sensors, labels, lineage)
Build and operate ETL/ingest pipelines that bring field recordings, synthetic data, and external deliveries into our cloud storage (GCP)
Read the full description (23 more sections)Show less
Own the data management and labeling lifecycle end-to-end: coordinate and communicate with external labeling companies and data subcontractors, track deliveries, run QA reports, and build the operational workflows they work in
Develop internal enabling tools for the whole AI organization: dataset search and filtering, APIs/backend, dashboards, and self-service data access
Run data migrations and indexing jobs; keep the catalog consistent and fast as data volume grows
Handle admin support and user access management — and then automate these support tasks so they stop being manual work
Establish good engineering hygiene in a young codebase: tests, typing, docs, logging, CI/CD
Shape the long-term architecture and vision of the data platform together with the team
Qualifications Strong Python
Solid SQL/PostgreSQL, including schema design
Experience with data modeling and metadata systems
Experience designing and operating ETL/data pipelines
Docker and CI/CD basics
Hands-on with object storage (GCS, S3, or similar)
Good software engineering hygiene: tests, docs, typing, logging
Organized and pragmatic: you can prioritize between a quick fix and a proper solution, and you know when each is right
Not allergic to support tasks — but technical enough to automate the support away
Comfortable coordinating with external vendors and non-technical stakeholders
Nice to have Familiarity with ML datasets and labeling workflows (images, video, lidar; annotation formats like COCO)
Experience with synthetic data generation or GenAI-assisted data workflows (auto-labeling, data augmentation, foundation-model-based curation)
Experience with GCP services beyond storage (BigQuery, Cloud Run, IAM)
Experience with data versioning / dataset tooling (DVC, LakeFS, FiftyOne, or similar)
Experience in a startup environment — comfortable with ambiguity and changing priorities
Exposure to robotics data formats (ROS bags, MCAP, PX4 logs)
Find more English Speaking Jobs in Germany on Arbeitnow
You will apply. Then you will hear nothing.
And no one will tell you what was wrong. See it before you send: your ATS score, every weak line, and the fix for each.
- 1
Drop in your CV
One PDF, thirty seconds. No card.
- 2
See what is wrong with it
Every weak passage, quoted from your own CV, with the line to replace it.
- 3
Apply where you fit
Every European role ranked against what your CV actually says.
What's actually stopping you?
- I apply and hear nothing backStart with the ATS scan — see what a filter does to your CV before a human sees it.Start here →
- I can't tell which roles I'd getStart with the match scores — every listing here ranked against what your CV actually says.Start here →
- Just browsing for nowKeep looking. Nothing to sign up for.
- Free, no card
- Your CV file is deleted after parsing
- Refreshed every 6 hours