ScoutsCapital
A social platform for football with a profile for everyone in the game: players, coaches, scouts, referees, clubs and leagues. Posts, video, live streaming and match results, so talent gets discovered through what it actually does.
- Role
- CTO
- When
- 2023 - now
- Links
- scoutscapital.com


The story
Tobias Schaller (Toby) came with a small idea: make football talent more discoverable. When we met, we decided to aim bigger. A social platform built around football, bringing together posts, video, live streaming and match results, with a profile for everyone in the game: players, coaches, scouts, referees, clubs and leagues. The experiences people know from Facebook, TikTok and Instagram, connected to football.
We wanted players to be discovered through what they actually do: their matches, footage and participation, rather than scores or claims they enter themselves.
Building that with a small team meant finding an architecture that could support our ambition at a fraction of the cost. I focused on making the infrastructure and streaming affordable enough to turn the idea into a working platform.
I have been CTO since day one. I lead a team of six developers, work alongside the CEO, CRO and CMO and their teams, and build hands-on: the platform together with a strong principal engineer who knows every layer of the stack, plus the backend, the web app and the moderation AI.
Our aim is to make football more discoverable and give talent a better chance to be seen.
Constraints
- A small team and a startup budget, building for iOS, Android and web at once.
- Traffic that spikes around matches, and fan uploads of photo and video at any hour.
- User-generated content that has to be moderated before anyone sees it.
- Payments and subscriptions across three app platforms.
Decisions I made, and why.
One API for every client
A single FastAPI and PostgreSQL backend serves web, iOS and Android, so features ship once and behave the same everywhere.
HLS streaming with three cache layers
My design: every video is transcoded into HLS so players fetch small segments at the right quality, then cached in three layers across Google Cloud Storage and Cloudflare. Video streaming costs dropped by 99%.
Kubernetes on GKE, not a PaaS
Video transcoding and ML inference need control over scaling and cost. Each workload gets its own autoscaling node pool, so a spike in one does not starve the others.
A job queue for heavy work
The API never transcodes itself. FastAPI queues jobs in Temporal, which spins up transcoding machines on demand and runs AI jobs that collect videos and posts by category for AI and analytics. Jobs survive restarts and retry on their own.
Preemptible machines where it is safe
Interruptible work like transcoding runs on preemptible nodes that scale up and down with demand, next to a small stable pool for the API and data. Scaling costs about 90% less than a traditional always-on setup.
Everything as code, two clusters
Terraform, Helm and ArgoCD rebuild any environment. A development cluster mirrors production with lighter protection, so the team can move fast without touching prod.
Locked down by default
Production is only reachable through bastion hosts, with RBAC, 2FA, App Check, rate limits and a dedicated firewall service in front of the API.
The trade-off
Kubernetes is more to run than a managed platform. I accepted that cost in exchange for control over scaling, spend and ML workloads, and paid it down with infrastructure as code, GitOps and full observability.
Problems we found, and fixed.
I led these calls and built the fixes together with our principal engineer, who knows every layer of the stack. Numbers come from the commits and measurements at the time.
A 15-second endpoint
- Problem
- The match line-up endpoint embedded about 18 full player profiles per row and timed out in the mobile app.
- Fix
- Fix: Dropped the data the client never read, without changing the API contract.
14.7 s to 0.03 s, 192 SQL queries to 2, payload down 57%.
Scaling on the right signal
- Problem
- The API autoscaled on memory rather than CPU, so it added capacity it did not need and slowed rollouts.
- Fix
- Fix: Moved scaling to CPU with tuned scale-up and scale-down behaviour, and tightened worker limits.
Leaner scaling and fast, predictable rollouts.
One owner for replica counts
- Problem
- ArgoCD and the autoscaler both adjusted replica counts, which caused needless pod churn.
- Fix
- Fix: Gave the autoscaler sole ownership of replicas and had ArgoCD ignore them.
About 1 s less latency per request.
A cheaper development cluster
- Problem
- The development environment ran far heavier than it needed to, monitoring included, around the clock.
- Fix
- Fix: Monitoring that scales to zero, smaller node pools and Redis, lighter logging, and idle IPs released.
About $1.5k a month saved.
Live video at the edge
- Problem
- Edge workers were placed next to the origin instead of the viewer, and stale manifests pushed live latency past 15 s.
- Fix
- Fix: Ran the worker at the viewer's edge, bypassed cache for live reloads and started playback close to the live point.
About 5x faster time to first byte.
Moderation that fails safe
- Problem
- Every upload needs a verdict, even when the cloud model is down or busy.
- Fix
- Fix: Frame checks on images and video keyframes, plus a dual audio-and-video model for violence. Batched requests to Vertex AI with a local fallback model, pay-per-use GPU jobs instead of an always-on endpoint, and a hash cache so identical media skips inference.
Every upload screened, with no idle GPU cost.
Sign-in and payment hardening
- Problem
- OAuth and one-time-password flows had account takeover paths, and payment webhooks could be processed twice.
- Fix
- Fix: Closed the takeover paths, added rate limits and stronger password hashing, and gave Stripe webhooks an idempotency ledger with a replayable dead-letter queue.
Takeover paths closed; each payment event handled once, failures replayable.
What I built
Leading teams and direction
I lead six developers and make the long-term technical calls: what we build next, how the platform evolves and where we invest.
Production Kubernetes
Designed with our principal engineer and run together on GKE across regions: workloads isolated on their own autoscaling node pools, sized automatically, and prioritised so critical services always get capacity.
Infrastructure as code and GitOps
Terraform and Helm for every environment, ArgoCD with automatic image updates, and builds on Cloud Build and GitHub Actions.
Content moderation AI, built in-house
I designed and built the moderation models myself in PyTorch: frame checks on images and video, and a dual audio-and-video model for violence. Served from Vertex AI inside the Temporal media pipeline.
Observability and security
Prometheus, Thanos, Loki, Grafana and Sentry in every service. Bastion-only access to production, TOTP 2FA, Firebase App Check, Turnstile and RBAC.
Edge and payments
Cloudflare Workers for storage, live and share images, Stripe and in-app purchases across web and mobile, and a Pub/Sub to BigQuery analytics pipeline.
Results
- Video streaming costs cut by 99% with HLS and three cache layers.
- Scaling costs about 90% lower than a traditional setup, with preemptible machines and autoscaling.
- Slowest API endpoint from 14.7 s to 0.03 s.
- Live on the App Store, Google Play and the web.
- Fan uploads are screened by our own moderation model.
- The whole platform is reproducible from code and deployed through GitOps.
Stack
- Nuxt 4
- Flutter
- FastAPI
- PostgreSQL
- Temporal
- GKE
- Terraform
- ArgoCD
- Cloudflare
The platform, part by part.
Apps
Edge
Services
AI
Data
Platform
Point at any part of the platform to see what I did there.
- Led and built hands-on
- Led
Dion