Data Engineer at Mediastream AG
I've been on ComplyGuard since its early months, in a team of about eight engineers and analysts. I started with a trial project through Upwork. Day to day I review and merge teammates' pull requests, and I've run sprint planning and stand-ups when our lead was away.
-
Two-layer tenant isolation on ClickHouse
A FastAPI query guard adds tenant filters from SSO group memberships. Row-level security policies, generated at deploy time, enforce the same scope inside ClickHouse, so a query that misses its tenant filter still can't return another tenant's rows. Fact tables without a tenant key are denied by default, and a 57-test isolation suite covers both layers.
-
Analytics API and a no-code dashboard builder
I built most of the analytics backend (now 100+ REST endpoints over 35 tables) and its first Svelte 5 front end. An earlier version of the customer-facing app, with the query guard and dashboard routes, runs in production. In the current version a decorator or a stored definition generates each dashboard's route, permission checks, tenant scoping and response model, and analysts build charts with live preview behind a SQL guard that rejects JOIN, UNION and CTE queries and requires the tenant filter.
-
Pipelines on Kubernetes
I deployed Airflow on Kubernetes for two data platforms. On a DNS-based web-filtering platform I split the URL-classification pipeline into independently scheduled scrape, translate and predict stages, so one failing stage no longer stops the others. I also built incremental Cassandra-to-ClickHouse paths (4 Kafka Connect connectors and a nightly PySpark job with high-watermarks), automated Feast feature refresh and NVIDIA Triton model deployment, and moved the MLflow tracking server onto Kubernetes.
-
Streaming KPIs from model scores
I built a streaming path from model scores to KPIs: Kafka Connect on Strimzi pulled AML and responsible-gambling predictions from Cassandra into Kafka, and a PyFlink job split them into 1-minute windowed KPIs and per-transaction score records in ClickHouse. A second PyFlink job rolled DNS hits up into daily per-domain KPIs. Both jobs were retired in 2026.
-
SSO, access control and audit
I set up Authelia with TOTP 2FA and LLDAP in front of the analytics app and Superset, with 9 data-driven role groups and GDPR-oriented audit logging. After a security review I changed the API to verify the session itself instead of trusting forwarded identity headers. I also found and closed a gap where a role limited to one dashboard area could fetch another area's data by URL.
-
Compliance workflows
With the team I built the rule-authoring side of a responsible-gambling rules engine (versioned rules, four-eyes approval, optimistic concurrency) and license, obligation and case management. State changes go through allow-listed state machines, and the audit history is written in the same transaction as the change.
-
Moves and reliability
I planned and verified the move of the production BI and Airflow stack to a new Kubernetes cluster: dashboards, hundreds of charts and about 20,000 historical DAG runs came across intact. Earlier I moved the analytics stack from Docker Compose to Kubernetes, set resource limits across 7 service deployments after OOM kills, and wrote a post-incident analysis of an intermittent cluster tunnel failure.