Follow Pay
From a simple Supabase backend to a distributed, real-time system.
Architecture evolution
- V1Supabase-first
Flutter UI, GetX, and Supabase — intentionally simple to ship the product idea quickly.
- V2FastAPI enters
As we thought about future load, Python/FastAPI took on more backend and API logic.
- Incident~10K traffic spike
Sudden traffic hit; Python servers went down. Approximately 14 hours of stabilization work followed.
- V3Supabase RPC shift
Moved Python-side functions into Supabase PostgreSQL RPC to reduce pressure on the Python servers.
- CostInfrastructure pressure
Several months later (~4–7 months in), Supabase and DigitalOcean bills had climbed too high.
- V4Self-hosted Supabase
Moved Supabase onto our own Docker-based infrastructure for cost and control.
- LaterTelegram verification in-app
Needed a Telegram-related verification flow inside Follow Pay. Existing Flutter/Telegram libraries were not reliable enough for the required path, so I investigated handy_tdlib and developed a working patch — used in production for about a year.
- V5Social + chat stack
Go, Dragonfly, WebSockets, REST, and local storage for social/chat responsiveness.
- V6State synchronization
Hardened WebSocket + REST + local-storage cooperation so chat state stayed consistent.
Starting point
Follow Pay started from the product idea. The first implementation was intentionally simple: a Flutter UI with GetX on top of a Supabase backend. At that stage, Supabase was enough to get the product moving quickly.
As the product grew, we started thinking about future load and introduced Python with FastAPI. FastAPI became responsible for handling more backend and API logic while the mobile client continued to evolve.
The ~10K traffic incident
At one point, approximately 10,000 traffic came in suddenly, and the Python servers went down. The architecture had been designed incrementally, but the spike exposed a weakness in how much of the system still depended on the Python layer staying healthy under pressure.
That night, we worked for approximately 14 hours to stabilize the system. The solution was not simply buying a bigger server. We moved Python-side functions into Supabase PostgreSQL RPC functions, reducing the amount of application logic that had to pass through the Python server.
It was an architectural response to a production failure: move critical work closer to the database path that could absorb the load more reliably in that moment.
Infrastructure cost pressure
Several months later — approximately 4–7 months into the project — another problem appeared. The system was working, but infrastructure costs had become too high. The Supabase bill increased. The DigitalOcean infrastructure bill increased.
At this stage I was increasingly working on the backend and infrastructure side and managing my own servers. I started asking whether we really needed managed infrastructure for everything.
Self-hosting Supabase
I decided to move Supabase from the managed environment onto our own infrastructure using Docker. The goal was greater control over infrastructure, deployment, cost, server resources, and backend architecture.
After moving Supabase to our own servers, the infrastructure became much smoother from a cost and control perspective. Exact percentage savings aren't claimed here — the meaningful change was ownership of the trade-offs.
Telegram verification: solving the in-app requirement
Follow Pay needed a Telegram-related verification flow, but sending users outside the app created limitations around the verification experience. The requirement was to keep that flow inside Follow Pay.
I evaluated multiple Flutter/Telegram libraries, but they did not reliably satisfy the required flow. I then investigated handy_tdlib more deeply — including how it talked to TDLib/libtdjson — and developed a patch/adaptation that allowed the integration to work for our case.
The patched implementation was used in Follow Pay in production for approximately one year. After maintaining that work and noticing continued developer demand for a Flutter TDLib integration around the original handy_tdlib package, I later published the maintained approach as handy_tdlib_next — a community-maintained package, not an official Telegram or TDLib release.
Social + chat architecture
Later the product needed social features and chat. That introduced a completely different class of engineering problems. The architecture used Go, Dragonfly, WebSockets, REST APIs, and local storage. Dragonfly became part of the real-time and chat infrastructure.
The goal was to make chat feel closer to the responsiveness users expect from applications such as WhatsApp — not just “messages exist,” but that the state feels immediate and trustworthy.
The real-time race-condition problem
We had multiple sources of state that could update the same conversation or message nearly simultaneously: WebSocket real-time events, REST API responses, and local storage for cached or persisted client state.
That created race conditions. In practice I encountered local messages being wiped unexpectedly, duplicated messages, incorrect ordering or state, WebSocket events arriving before or after REST responses, and local state drifting away from server state.
The hard part was not making WebSockets work. The hard part was making multiple asynchronous sources of truth cooperate without corrupting local state.
Real-time systems are rarely difficult because of the WebSocket itself. They become difficult when multiple asynchronous systems believe they can update the same state.
How I worked through it
I had to reason about event ordering, local persistence, REST synchronization, WebSocket events, duplicate events, state reconciliation, race conditions, and the full message lifecycle.
After repeatedly reproducing and debugging the inconsistent states, I redesigned the synchronization behavior until the local state and real-time state behaved consistently. The work was less about a single clever trick and more about making the entire pipeline honest about who could mutate what, and when.
Technical architecture
Architecture evolved into multiple services and state layers
Not every request follows this exact linear path — this is the evolved service and state layering.
- Flutter + GetX
Client application and state management
- Local state / storage
Isar · SQLite · Hive
- API layer
FastAPI
- Backend services
Go
- PostgreSQL / Supabase
Data + RPC paths (including self-hosted Docker)
- Real-time infrastructure
Dragonfly + WebSockets
- Telegram / TDLib path
In-app verification via patched handy_tdlib work — later published as handy_tdlib_next
What this taught me
- Start simple, but design with the next bottleneck in mind.
- Production traffic reveals architectural assumptions.
- Managed infrastructure is convenient, but infrastructure ownership changes the trade-offs.
- When libraries do not fit a production requirement, reading the integration layer carefully can unlock a workable path.
- Real-time applications require explicit state synchronization strategies.
- Debugging race conditions requires understanding the entire state pipeline, not just one component.