← All work

JAVIS Technologies

Async Execution Framework

A Spring Boot starter that gave every backend service correctness by default

JavaSpring BootMySQLResilience4j
Problem
Every team was writing its own async plumbing — job lifecycle, retries, idempotency, completion callbacks. Each implementation was subtly different, each had its own crash-window bugs, and correctness depended on whichever engineer wrote it that quarter.
My role
Designed the library and drove adoption across all backend services.
Approach
Pushed the hard guarantees into a platform library rather than a set of guidelines, so domain teams implement only a business handler and get the rest structurally. Effectively-exactly-once falls out of at-least-once delivery plus idempotent effects: idempotency keys, compare-and-set state transitions, lease/heartbeat/reaper crash recovery, and zombie-write protection. The constraint is that handlers must be written to be replayable, which is a real thing to ask of every team.
Outcome
Adopted by every backend service in the org. Async boilerplate disappeared from domain code, and the crash-recovery guarantees are validated by CI chaos tests across all three dual-write crash windows.

Why a library and not a guideline

The failure mode of “everyone implements async correctly” at organizational scale, and what convinced you it had to be structural.

The correctness argument

At-least-once delivery plus idempotent effects, and why that composition gives you effectively-exactly-once. Worth walking through one concrete interleaving — this is the part an interviewer will want to dig into.

Crash recovery

Leases, heartbeats, the reaper. What happens to a job whose worker dies holding it, and how zombie writes are prevented when it comes back.

Proving it

The CI chaos tests: the three dual-write crash windows, how you inject the faults, and what the tests assert. This section is the one that separates a claim from evidence.