Event Sourcing
This post is about event sourcing and why the fundamentals of event sourcing matter in this age of agentic software.
Before I introduce event sourcing, let me tell you who I am. I have worked in financial industry for about half-a-decade, and I recently build HFT infrastructure from scratch. While building that, I found certain patterns that are helpful for me, and also helpful for agents, to debug faster, to reproduce the issue exactly, and to verify the fixes.
Event sourcing123 is about journaling the events which cause transitions in your state machine. What is a state machine, you may ask? A state machine is a software's internal status which determines the next business outcome.
Why does event sourcing help to debug a bug faster? Event sourcing is an architectural decision which allows you to replay the software's path in an offline manner. Whenever we are stuck reproducing what might have happened by looking at the logs, the network packets, or different microservices, a simple event sourcing system lays out in bare bones the events that took place, which caused the software to end up in the state that it did. The bug reproduction becomes very fast, and that is one of the ideal goals of an event-sourced system.
Why event logs matter#
So let me explain it a bit more clearly: what exactly is an event source system, and where have we been seeing one in practice? Let's say you have a bank account. Your bank account is a summation of all the credits and debits of your money. Now imagine one day you go up to your bank, and it says that you have to spend dollars. When you checked last night, you had 100k. So you will demand the transaction ledger for this, right? You are not ready to accept the bank's website pixels displaying that you have 10 dollars.
An event source system is kind of like that. It has a transactional ledger which keeps account of all the state changes, all the state mutations that lead up to your current state. In a way, it makes the state of an application auditable.
Now, since I have worked in the finance industry, I see its applications everywhere. It helps to reproduce bugs, to actually see what events preceded the bug happening, and so on.
Why deterministic replay matters#
Starting from a snapshot, if you keep replaying the events that are in the journal, you should end up with the current state. That's the ultimate goal for an event-sourced system. What makes it challenging are certain bits of non-determinism sprinkled through software. Non-determinism comes from the kernel. It can come from network jitter. It can come from the fact that we are not modeling our events properly.
Having the ability to replay deterministically what changes to the state happened in production is one of the key guarantees of an event-sourced system, and that is what makes the paradigm so powerful.
Challenges of replaying event logs#
These challenges all come down to non-determinism. Event sourcing has several such sources.
Hidden inputs. Can an event-sourced system replay everything? Replay every bug? How does it know what the environment variables were? Which feature flags was the software using? If these aren't captured, the same events can produce different behavior on replay. As a company, Kavach Labs is going to tackle these challenges so that replay becomes truly deterministic.
Multiple views and replicated state. Software often keeps state in more than one place. Say you design an event store that writes to an in-memory database. How do you replay without knowing the state of that database? With 1,000 rows it may behave one way, but on replay you are working with a fresh database clone.
State that isn't derived from the journal. One system requirement for replayability is that your state should be a first-order derivative of the journal. Most single-threaded deterministic software follows this principle, but quite a few chunks of software still don't. How do you model those? These are the uncomfortable parts of the system that we want to tackle.
We're building Kavach Labs to tackle this problem head-on. We love event sourcing as an architectural concept, and we want agentic software to be written and deployed so that it's easy to capture a bug, replay it, fix the code, and deploy the fix with more confidence.
References
- ArticleEvent Sourcingmartinfowler.com↩
- VideoEvent Sourcing - Martin Fowler (YOW! 2016)classcentral.com↩
- ArticleMy notes on the talk "The Many Meanings of Event-Driven Architecture" by Martin Fowler (GOTO 2017)gist.github.com↩