From a product problem to a reusable platform.
We didn’t set out to build a backend platform. When we started developing ExME, our priority was simple: move quickly. We needed the usual backend capabilities — authentication, database access, file handling, email delivery and APIs — without spending months building infrastructure before we had a product.
At the time, DreamFactory was a good fit. It was open source, had no licensing cost, and gave us many of the capabilities we needed out of the box. For an early-stage product, that trade-off made sense. Then ExME grew. And the assumptions behind that decision started to change.
When the problem changes, the architecture has to change
As ExME grew in users and complexity, we started to encounter stability and scalability problems. The issue wasn’t simply that the system was “slow”. The more relevant problem was resource efficiency. We were seeing a relatively high consumption of resources for the amount of requests being processed. The architecture that had been perfectly adequate when the product was smaller was becoming increasingly difficult to scale efficiently.
At the same time, our requirements around data security were becoming more demanding. In particular, implementing the level of row-level security and data access control we wanted required increasingly complex workarounds. Then there was another change: the commercial model of the technology no longer matched the reason we had originally chosen it.
At that point, continuing with the same solution meant accepting compromises in several areas simultaneously:
- scalability;
- resource consumption;
- security and data access control;
- operational complexity;
- and cost.
That was the point at which we stopped asking:
“How can we make this solution work?”
and started asking:
“What do we actually need from our backend?”
Building only what we needed
The answer wasn’t to build another generic backend framework. We wanted something focused on the requirements we were actually dealing with:
- authentication and authorisation;
- controlled access to data;
- predictable performance;
- file management;
- email communication;
- structured logging;
- configuration;
- scalability;
- and, importantly, consistency across applications.
We built the new backend using Spring Boot, together with a configuration portal that allows much of the platform behaviour to be managed without repeatedly implementing the same functionality in every application. The goal wasn’t to create another technology stack for the sake of having one. The goal was to centralise capabilities that were becoming common across our applications.
From ExME backend to Fabric
Initially, the new backend existed because ExME needed it. That changed relatively quickly. As we developed other applications, we kept encountering the same requirements:
- How should users authenticate?
- How should an application access a particular piece of data?
- How should files be stored?
- How should operations involving several backend actions be executed?
- How should failures be logged?
- How should applications be configured consistently?
We already had solutions for many of these problems. There was little reason to solve them independently again. So the backend stopped being the backend for ExME and became something more general: a reusable platform for our software projects.
That distinction matters. Reusing code is useful. Reusing well-understood architectural decisions is much more valuable.
The platform is becoming an ecosystem
The Fabric is evolving into a set of specialised components rather than one large application responsible for everything. At the moment, three components define the direction of that architecture.
Fabric
│
┌──────────────────┼──────────────────┐
│ │ │
Heimdall Saga Ekko
Authentication Operations Communication
│ │ │
Identity Synchronous Real-time
Security execution + Push
│
Multiple steps
Each component has a specific responsibility.
Heimdall — Identity and authentication
Heimdall is responsible for the identity and authentication side of the ecosystem. It provides the common foundation for applications that need to know:
- who the user is;
- whether the user is authenticated;
- what access is available;
- and how that identity is propagated through the backend.
This is an important distinction for us. Authentication shouldn’t be something that every application implements independently. It is infrastructure.
Saga — executing an operation
The component currently known internally as Pilot has also evolved beyond the idea of a simple worker. Its job is to execute operations composed of multiple steps.
For example, consider: Send an email to user 2. That operation might contain:
Operation: sendEmail
Step 1
→ Database
→ Get email address for user 2
Step 2
→ Email
→ Send message
Another operation might be: Update the avatar of user 4.
Operation: uploadFileForAvatar
Step 1
→ Storage
→ Upload file
Step 2
→ Database
→ Store the resulting file reference
The individual steps can have different types and interact with different backend capabilities. This gives us a common execution model instead of implementing these multi-stage operations separately inside every application. We are currently using this model for synchronous operations. The architecture also leaves room for a future asynchronous execution capability, which becomes increasingly relevant as the number and complexity of operations grow.
Ekko — communicating with users
Ekko is the next part of the ecosystem we are developing. Its responsibility is communication between our applications and their users, particularly where communication needs to happen in real time.
The intended model is not simply: “Send a notification.” It is closer to:
Application
│
▼
Ekko
│
├── Real-time communication
│
└── Push notification fallback
If a user is actively connected, communication can happen through the real-time channel. If they aren’t, the system can fall back to push notifications. This creates another capability that applications shouldn’t need to implement independently.
Where the platform is going
We don’t see the Fabric as a finished product. In fact, one of the useful characteristics of the platform is that its roadmap is being driven by problems we actually encounter. The direction currently includes areas such as:
Authentication and authorisation
Continue evolving the identity and access-control model across applications.
Operation execution
Expand the range of available steps and improve execution control, error handling and observability.
Real-time communication
Complete Ekko and its integration with push notification fallback.
Asynchronous processing
Introduce background execution for operations that shouldn’t remain tied to a synchronous request. That will introduce another set of concerns:
- queues;
- retries;
- failure handling;
- idempotency;
- execution state;
- monitoring.
Those aren’t implementation details we can simply add later without thinking about them. They are part of the architecture.
Observability
As the platform becomes shared infrastructure, understanding what it is doing becomes increasingly important. Logging is only the starting point. We need to be able to understand:
- what happened;
- when it happened;
- where it happened;
- what failed;
- and what happened afterwards.
Why build this ourselves?
This is probably the question we should ask explicitly. The lesson isn’t:
“Build instead of buy.”
We don’t believe that. The original decision to use an existing backend was reasonable. Building authentication, file management, email delivery and APIs from scratch before ExME had validated its own requirements would not necessarily have been a better engineering decision.
The lesson was different. The right architecture can change as the problem changes. At the beginning, using an existing platform gave us speed. Later, the same platform introduced constraints around scalability, resource consumption, security and cost.
Building our own backend became justified because we were no longer solving one isolated problem. We were solving the same class of problems across several applications. That changed the economics and the architecture.
The platform is not the product
There is another principle behind the Fabric that is perhaps more important than the technology itself. We don’t want applications to exist simply to demonstrate the platform. The opposite is true. The platform exists to make applications easier to build, operate and evolve.
That means we don’t want to dictate technology because we happen to have built something internally. If a different technology is better for a particular problem, we should use it. The platform should remove unnecessary complexity, not create a new form of it.
From one problem to a platform
The evolution can therefore be summarised quite simply:
ExME
│
│ Need to move quickly
▼
Existing backend platform
│
│ ExME grows
│
├── Scalability constraints
├── Resource consumption
├── Security requirements
├── Data access complexity
└── Commercial model changes
│
▼
Fabric
│
├── Heimdall
│ Identity & Authentication
│
├── Saga
│ Operation execution
│
└── Ekko
Real-time communication
+ Push fallback
│
▼
Reusable platform
│
└── Multiple applications
We didn’t start with a plan to build a backend platform. We started by trying to solve a problem. The platform emerged when we realised that the problem wasn’t unique to ExME.
And that is probably the most important architectural lesson in this story: Good architecture isn’t about making the right decision once. It’s about recognising when the problem has changed — and being willing to change with it.