← Back to Featured WorkEngineering Case Studies

How the systems evolved.

The interesting part of building software isn't only what ships. It's what happens when the first architecture meets real traffic, real-time state, unreliable networks, and production constraints.

Case Study 01

Follow Pay

From a simple Supabase backend to a distributed, real-time system.

FlutterGetXIsarSQLiteHiveFastAPIGoPostgreSQLSupabaseDragonflyWebSockets

Architecture evolution

  1. V1
    Supabase-first

    Flutter UI, GetX, and Supabase — intentionally simple to ship the product idea quickly.

  2. V2
    FastAPI enters

    As we thought about future load, Python/FastAPI took on more backend and API logic.

  3. Incident
    ~10K traffic spike

    Sudden traffic hit; Python servers went down. Approximately 14 hours of stabilization work followed.

  4. V3
    Supabase RPC shift

    Moved Python-side functions into Supabase PostgreSQL RPC to reduce pressure on the Python servers.

  5. Cost
    Infrastructure pressure

    Several months later (~4–7 months in), Supabase and DigitalOcean bills had climbed too high.

  6. V4
    Self-hosted Supabase

    Moved Supabase onto our own Docker-based infrastructure for cost and control.

  7. Later
    Telegram verification in-app

    Needed a Telegram-related verification flow inside Follow Pay. Existing Flutter/Telegram libraries were not reliable enough for the required path, so I investigated handy_tdlib and developed a working patch — used in production for about a year.

  8. V5
    Social + chat stack

    Go, Dragonfly, WebSockets, REST, and local storage for social/chat responsiveness.

  9. V6
    State synchronization

    Hardened WebSocket + REST + local-storage cooperation so chat state stayed consistent.

Starting point

Follow Pay started from the product idea. The first implementation was intentionally simple: a Flutter UI with GetX on top of a Supabase backend. At that stage, Supabase was enough to get the product moving quickly.

As the product grew, we started thinking about future load and introduced Python with FastAPI. FastAPI became responsible for handling more backend and API logic while the mobile client continued to evolve.

The ~10K traffic incident

At one point, approximately 10,000 traffic came in suddenly, and the Python servers went down. The architecture had been designed incrementally, but the spike exposed a weakness in how much of the system still depended on the Python layer staying healthy under pressure.

That night, we worked for approximately 14 hours to stabilize the system. The solution was not simply buying a bigger server. We moved Python-side functions into Supabase PostgreSQL RPC functions, reducing the amount of application logic that had to pass through the Python server.

It was an architectural response to a production failure: move critical work closer to the database path that could absorb the load more reliably in that moment.

Infrastructure cost pressure

Several months later — approximately 4–7 months into the project — another problem appeared. The system was working, but infrastructure costs had become too high. The Supabase bill increased. The DigitalOcean infrastructure bill increased.

At this stage I was increasingly working on the backend and infrastructure side and managing my own servers. I started asking whether we really needed managed infrastructure for everything.

Self-hosting Supabase

I decided to move Supabase from the managed environment onto our own infrastructure using Docker. The goal was greater control over infrastructure, deployment, cost, server resources, and backend architecture.

After moving Supabase to our own servers, the infrastructure became much smoother from a cost and control perspective. Exact percentage savings aren't claimed here — the meaningful change was ownership of the trade-offs.

Telegram verification: solving the in-app requirement

Follow Pay needed a Telegram-related verification flow, but sending users outside the app created limitations around the verification experience. The requirement was to keep that flow inside Follow Pay.

I evaluated multiple Flutter/Telegram libraries, but they did not reliably satisfy the required flow. I then investigated handy_tdlib more deeply — including how it talked to TDLib/libtdjson — and developed a patch/adaptation that allowed the integration to work for our case.

The patched implementation was used in Follow Pay in production for approximately one year. After maintaining that work and noticing continued developer demand for a Flutter TDLib integration around the original handy_tdlib package, I later published the maintained approach as handy_tdlib_next — a community-maintained package, not an official Telegram or TDLib release.

Social + chat architecture

Later the product needed social features and chat. That introduced a completely different class of engineering problems. The architecture used Go, Dragonfly, WebSockets, REST APIs, and local storage. Dragonfly became part of the real-time and chat infrastructure.

The goal was to make chat feel closer to the responsiveness users expect from applications such as WhatsApp — not just “messages exist,” but that the state feels immediate and trustworthy.

The real-time race-condition problem

We had multiple sources of state that could update the same conversation or message nearly simultaneously: WebSocket real-time events, REST API responses, and local storage for cached or persisted client state.

That created race conditions. In practice I encountered local messages being wiped unexpectedly, duplicated messages, incorrect ordering or state, WebSocket events arriving before or after REST responses, and local state drifting away from server state.

The hard part was not making WebSockets work. The hard part was making multiple asynchronous sources of truth cooperate without corrupting local state.

Real-time systems are rarely difficult because of the WebSocket itself. They become difficult when multiple asynchronous systems believe they can update the same state.

How I worked through it

I had to reason about event ordering, local persistence, REST synchronization, WebSocket events, duplicate events, state reconciliation, race conditions, and the full message lifecycle.

After repeatedly reproducing and debugging the inconsistent states, I redesigned the synchronization behavior until the local state and real-time state behaved consistently. The work was less about a single clever trick and more about making the entire pipeline honest about who could mutate what, and when.

Technical architecture

Architecture evolved into multiple services and state layers

Not every request follows this exact linear path — this is the evolved service and state layering.

  1. Flutter + GetX

    Client application and state management

  2. Local state / storage

    Isar · SQLite · Hive

  3. API layer

    FastAPI

  4. Backend services

    Go

  5. PostgreSQL / Supabase

    Data + RPC paths (including self-hosted Docker)

  6. Real-time infrastructure

    Dragonfly + WebSockets

  7. Telegram / TDLib path

    In-app verification via patched handy_tdlib work — later published as handy_tdlib_next

What this taught me

  • Start simple, but design with the next bottleneck in mind.
  • Production traffic reveals architectural assumptions.
  • Managed infrastructure is convenient, but infrastructure ownership changes the trade-offs.
  • When libraries do not fit a production requirement, reading the integration layer carefully can unlock a workable path.
  • Real-time applications require explicit state synchronization strategies.
  • Debugging race conditions requires understanding the entire state pipeline, not just one component.
Case Study 02

Alo Voice

When calling systems stop behaving like ordinary mobile applications.

React NativeNode.jsTelnyxKotlin / Native AndroidCallKeepVoIP / SIP

Starting point

Alo Voice is a React Native calling application with a Node.js backend and Telnyx for calling infrastructure, plus substantial native Android work and CallKeep integration.

The product needed to behave like a proper phone or calling application — not simply fire an API call and hope the UI catches up. That meant coordinating React Native, native Android, CallKeep, the SIP/VoIP layer, and Telnyx-backed calling infrastructure.

Race-condition classes we hit

During development and debugging I encountered race-condition classes around call readiness and event ordering. These were not claimed as every-call failures — they were intermittent, timing-sensitive failure modes that had to be hunted down.

  • A call could arrive before CallKeep was ready.
  • SIP could receive a call while another part of the application was still initializing.
  • Call UI state and SIP state could become temporarily inconsistent.
  • The native call lifecycle could move faster than the React Native layer.
  • A call could technically arrive, but the expected CallKeep presentation would not happen.
  • Initialization order could change whether the call was handled correctly.

Why this was difficult

This was not simply a React Native bug. Multiple asynchronous systems with independent lifecycles were involved: the React Native runtime, Android process/lifecycle, SIP signaling, CallKeep, VoIP events, and network availability.

The challenge was determining which system was ready, which event happened first, and where the state transition was being lost between layers.

Engineering approach

I worked through the lifecycle and race conditions by coordinating initialization and event handling across the native and React Native layers — native Android architecture, CallKeep integration, VoIP lifecycle, event ordering, and native ↔ React Native communication.

The useful outcome was not a single patch. It was a clearer model of how call state must move through the stack without assuming every layer is ready at the same time.

In VoIP, timing is part of correctness.

Technical architecture

Calling stack across multiple lifecycles

Not every request follows this exact linear path — this is the evolved service and state layering.

  1. React Native

    Application UI and JS runtime

  2. Native Android

    Kotlin / platform call handling

  3. CallKeep

    Native call UI / system presentation

  4. SIP / VoIP layer

    Signaling and call media path

  5. Telnyx / calling infrastructure

    External calling infrastructure

  6. Node.js backend

    Application backend services

What this taught me

  • Calling products inherit the complexity of every layer they touch.
  • Independent lifecycles create timing bugs that look like UI bugs.
  • Native ↔ React Native communication has to be designed, not improvised.
  • Reproducing race conditions is part of the engineering work.
Case Study 03

BEMXC

Designing a trading system around high-volume market data.

Kotlin MultiplatformRustMT5 environmentHigh-frequency trading design

Architecture

BEMXC is a trading platform built around a Kotlin Multiplatform client and a Rust backend. It was designed as a high-frequency trading platform — the architecture had to respect continuous market data and performance-sensitive server paths from the start.

The architecture separated the MT5 environment from the Rust service responsible for distributing critical trading data, allowing the backend workload to be handled independently.

The scaling problem

The challenge was not simply building CRUD APIs. A trading-oriented system has to deal with continuous data, many consumers, rapid updates, server load, efficient serialization and transport, backend resource management, and reliability under load.

Rust was chosen for the performance and control requirements of this architecture. No specific latency or throughput benchmark is claimed here — the point is the design pressure those workloads create.

Engineering direction

BEMXC connects to the broader portfolio story: Kotlin Multiplatform on the client, Rust on the backend, performance-oriented architecture, high-frequency trading requirements, separation of workloads, and server-side scalability concerns.

Different systems demand different abstractions. For BEMXC, the architecture had to be designed around the data flow first.

Technical architecture

Workload separation for trading data

Not every request follows this exact linear path — this is the evolved service and state layering.

  1. Kotlin Multiplatform client

    Shared client architecture for the trading product

  2. Rust service

    Distributing / processing critical trading-related data

  3. MT5 environment

    Dedicated main server/environment for MT5

  4. Consumers

    Large numbers of consumers receiving trading data

What this taught me

  • Trading systems force you to design around data flow, not screens.
  • Separating workloads keeps critical paths from competing with unrelated server work.
  • Performance-oriented backends are a product decision, not only a language preference.