Portrait of Ceyhun Akar

Ceyhun Akar

Data engineer. I build the secure, multi-tenant analytics layer and the pipelines that feed it.

I collect data, process it, and carry it to the point where a decision gets made. Most of my recent work makes the safe path the default: the API adds the tenant filter for you, nobody approves their own rule, and nobody can edit the history.

Work

private code, public description
Dec 2024 - nowContract, remote

Data Engineer at Mediastream AG

I've been on ComplyGuard since its early months, in a team of about eight engineers and analysts. I started with a trial project through Upwork. Day to day I review and merge teammates' pull requests, and I've run sprint planning and stand-ups when our lead was away.

  • Two-layer tenant isolation on ClickHouse

    A FastAPI query guard adds tenant filters from SSO group memberships. Row-level security policies, generated at deploy time, enforce the same scope inside ClickHouse, so a query that misses its tenant filter still can't return another tenant's rows. Fact tables without a tenant key are denied by default, and a 57-test isolation suite covers both layers.

    • ClickHouse
    • FastAPI
    • Authelia
    • LLDAP
    • pytest
  • Analytics API and a no-code dashboard builder

    I built most of the analytics backend (now 100+ REST endpoints over 35 tables) and its first Svelte 5 front end. An earlier version of the customer-facing app, with the query guard and dashboard routes, runs in production. In the current version a decorator or a stored definition generates each dashboard's route, permission checks, tenant scoping and response model, and analysts build charts with live preview behind a SQL guard that rejects JOIN, UNION and CTE queries and requires the tenant filter.

    • FastAPI
    • Pydantic
    • SQLAlchemy
    • Redis
    • Svelte 5
    • ECharts
  • Pipelines on Kubernetes

    I deployed Airflow on Kubernetes for two data platforms. On a DNS-based web-filtering platform I split the URL-classification pipeline into independently scheduled scrape, translate and predict stages, so one failing stage no longer stops the others. I also built incremental Cassandra-to-ClickHouse paths (4 Kafka Connect connectors and a nightly PySpark job with high-watermarks), automated Feast feature refresh and NVIDIA Triton model deployment, and moved the MLflow tracking server onto Kubernetes.

    • Airflow
    • KubernetesPodOperator
    • PySpark
    • Kafka Connect
    • Helm
    • MLflow
  • Streaming KPIs from model scores

    I built a streaming path from model scores to KPIs: Kafka Connect on Strimzi pulled AML and responsible-gambling predictions from Cassandra into Kafka, and a PyFlink job split them into 1-minute windowed KPIs and per-transaction score records in ClickHouse. A second PyFlink job rolled DNS hits up into daily per-domain KPIs. Both jobs were retired in 2026.

    • Kafka
    • Kafka Connect
    • PyFlink
    • Cassandra
    • ClickHouse
  • SSO, access control and audit

    I set up Authelia with TOTP 2FA and LLDAP in front of the analytics app and Superset, with 9 data-driven role groups and GDPR-oriented audit logging. After a security review I changed the API to verify the session itself instead of trusting forwarded identity headers. I also found and closed a gap where a role limited to one dashboard area could fetch another area's data by URL.

    • Authelia
    • LLDAP
    • Superset
    • PostgreSQL
  • Compliance workflows

    With the team I built the rule-authoring side of a responsible-gambling rules engine (versioned rules, four-eyes approval, optimistic concurrency) and license, obligation and case management. State changes go through allow-listed state machines, and the audit history is written in the same transaction as the change.

    • FastAPI
    • SQLAlchemy
    • PostgreSQL
    • React
    • TanStack
  • Moves and reliability

    I planned and verified the move of the production BI and Airflow stack to a new Kubernetes cluster: dashboards, hundreds of charts and about 20,000 historical DAG runs came across intact. Earlier I moved the analytics stack from Docker Compose to Kubernetes, set resource limits across 7 service deployments after OOM kills, and wrote a post-incident analysis of an intermittent cluster tunnel failure.

    • Kubernetes
    • Rancher
    • Superset
    • Airflow

Stack

honest tiers

Work stack used in paid work or a live product Built with hands-on in my own projects

Streaming
Apache Kafka Kafka Connect Apache Flink (PyFlink) Spark Structured Streaming Confluent Cloud
Batch and orchestration
Apache Airflow Apache Spark (PySpark) dbt Apache NiFi
Storage and query
ClickHouse PostgreSQL Apache Cassandra Redis MySQL Elasticsearch Trino MongoDB Apache Hive MinIO
BI and visualization
Apache Superset Apache ECharts Kibana Tableau
Platform
Kubernetes Helm Rancher (RKE2) k3s Docker AWS (S3, ECR, Translate) Google Cloud Authelia and LLDAP NGINX Ingress GitHub Actions
Apps and APIs
FastAPI Pydantic SQLAlchemy pytest Svelte React Next.js
ML and AI
MLflow Feast NVIDIA Triton Hugging Face Transformers PyTorch Amazon Bedrock CrewAI
Observability
Prometheus Grafana Loki Logstash (ELK)
Languages
Python SQL Bash TypeScript
Web scraping
Playwright aiohttp
Trained
Databricks Amazon Redshift AWS Glue Amazon Athena Apache Hadoop Kafka Streams ksqlDB IBM Cognos Analytics Couchbase
Exploring
Terraform Apache Beam Dagster Neo4j Looker Studio

Projects

2024, code with its video and write-up

pipeline · 23 stars

Streaming API records into Cassandra

Airflow pulls user records from an API into Kafka, and Spark Structured Streaming processes them into Cassandra. Everything runs in Docker Compose.

  • Airflow
  • Kafka
  • Spark
  • Cassandra
  • Docker

pipeline · 5 stars

Reddit data platform

Reddit API data moves through NiFi and Airflow into MinIO and Hive, is queried with Trino, and ends up in Superset and Tableau dashboards.

  • NiFi
  • Airflow
  • MinIO
  • Hive
  • Trino
  • Superset

pipeline · 2 stars

One-time code delivery pipeline

Airflow writes email addresses to a three-broker Kafka cluster, loads them into Cassandra and MongoDB, checks the email and one-time code in both, and sends the code by email, Slack and Discord.

  • Airflow
  • Kafka
  • Cassandra
  • MongoDB
  • Docker

I built these in 2024 to learn the tools end to end. My work at Mediastream lives in private repositories, and the Work section above describes it.

Education

  • Executive MBA (non-thesis, evening program), Istanbul Technical University2026 - now
  • B.Sc. Metallurgical and Materials Engineering, Istanbul Technical University
  • Udacity Data Engineering Nanodegree, IBM Data Engineering Specialization, Apache Kafka and Elasticsearch (BTK Academy)2023 - 2024

Contact

Email is the fastest way to reach me, whether it's about a role, a contract or something I wrote.