A Slack message has two lives. First, it is part of a conversation happening now. Later, it becomes something a colleague searches for: a decision, a deployment note or the explanation behind an old project.

Supporting both lives creates an interesting storage problem. The system must save messages reliably and deliver them quickly, but also retrieve useful answers from years of history without exposing conversations the searcher cannot access.

Introduction

Slack's engineering publications describe MySQL storage scaled through Vitess and a separate search system built around Solr and Lucene. These published accounts document particular stages of the architecture, rather than every component running at Slack today.

We will use those facts as a foundation, then follow a fictional team's project channel to explain the underlying decisions. Example schemas and implementation suggestions are teaching models. The emphasis is on why storage, live delivery and search need different responsibilities even though the user experiences one application.

Durable Messages Come Before Live Delivery

Slack's 2020 datastore article explains that messages are persisted before they are distributed through its real-time WebSocket infrastructure. The same article describes its migration from earlier MySQL arrangements to Vitess, a system for horizontally scaling MySQL. Slack's Vitess architecture

This ordering makes the message history a durable foundation. A colleague whose laptop is asleep can later retrieve the message even though they missed the live event.

Imagine Priya writes “Release approved” in project-orbit. A simplified service validates her membership and stores the message before confirming success. Connected clients can then receive an event telling them something changed.

The event is a delivery mechanism, while the stored record is the lasting conversation. Confusing those roles leads to systems that work beautifully while everyone is connected but lose context whenever a device reconnects.

Choosing a Partition Key Changes the Product's Limits

Slack's published migration account describes moving beyond workspace-based sharding, including the motivation to distribute message data by channel ID. A very large workspace should not force all its message traffic onto one shard.

To understand the tradeoff, imagine storing every record for a customer on one database. Many customer-specific operations become easy, and isolation is straightforward. But the largest customer eventually becomes the largest indivisible unit of load.

Splitting by channel gives more flexibility. Different channels in the same workspace can live in different places. In return, an operation covering the entire workspace may need to contact several partitions.

There is no universally perfect key. The useful questions are which operations happen most often, which groups can grow without bound, and which requests would become expensive after partitioning. The partition key should reflect those answers rather than merely matching the top-level entity in the interface.

A Message Has More Structure Than Its Text

A teaching schema for project-orbit might include a message ID, channel ID, sender ID, creation time, body, thread reference and edit version. Attachments can be referenced separately so ordinary history reads remain small.

Each field supports a behaviour. The channel identifies the conversation. The ordering fields support pagination. The thread reference connects replies. The version helps distinguish a recent edit from an older copy.

It is tempting to encode the entire channel as one large document. That makes loading a small channel easy, but every edit then competes to modify an increasingly large object. Separate message records allow bounded reads and more focused updates.

The UI can still present one continuous timeline. A smooth interface does not require one giant storage record; it requires an API that assembles the right bounded slice of records consistently.

Search Is a Separate Representation

Slack's search engineering article describes Solr, built on Lucene, and the use of ranking signals to improve the usefulness of results. A text index serves a different access pattern from a chronological message table. Search at Slack

If someone searches for “release approved,” scanning every stored message would waste work. An inverted index associates searchable terms with the records containing them, narrowing the candidate set before ranking.

The searchable document may include more than the message text. In a conceptual design, channel, author and time become structured fields supporting filters. Those fields should have deliberate types: a timestamp is more useful as a sortable date than as an arbitrary string.

The index is derived data. It needs to follow edits and removals, and the team needs a recovery strategy if an index must be rebuilt. Treating it as an independent, uncoordinated database creates conflicting versions of the conversation.

Following a Message into Search

Consider a possible implementation for our project channel. After the message is committed, a durable change event tells an indexing worker to update the search representation. The worker transforms the message into searchable fields and acknowledges the event after successful processing.

This decouples search availability from message sending. If indexing pauses briefly, colleagues can still talk, although their newest messages may not appear in search immediately.

The delay needs an observable measure. “Search is healthy” should include how old the newest indexed data is, not simply whether the search endpoint responds. A fast answer from yesterday's index may be operationally successful but useless to someone looking for today's decision.

Retries also need ordering rules. If version three of a message reaches the index before version two, the older event must not replace it. The exact mechanism can vary; the requirement is that eventual processing converges to the newest valid state.

Permissions Belong in the Search Path

A search engine can find text without understanding who is allowed to read it. That is dangerous in a workplace containing private channels and conversations across different groups.

In our design, search must restrict results according to the searcher's current access. It should apply that restriction before exposing message snippets, counts or highlights, not merely hide the full message after a click.

Suppose someone leaves a private project channel. An old cache of their memberships could allow a search result to reveal information they should no longer see. Permission data therefore has its own freshness requirements, distinct from whether an ordinary message edit can take a moment to appear.

Slack's later security discussion for enterprise search reinforces the importance of respecting source permissions, although connected enterprise content has additional concerns beyond channel history. Slack's enterprise search security design

Threads Change Retrieval Without Replacing the Timeline

A threaded reply belongs to a conversation but also to a smaller discussion inside it. That creates at least two useful reads: the main channel history and the replies beneath a particular parent message.

Our schema could represent a reply with a root-message reference and an ordering key. Opening a thread then retrieves its replies directly rather than scanning the entire channel and filtering in application memory.

Deleting or editing the root introduces edge cases. Does the thread remain visible with a placeholder? Are replies still searchable? Those are product decisions that must be reflected consistently in storage and indexing.

The important engineering lesson is that a user-visible relationship often needs an explicit query path. If threads are a central feature, they should not be an expensive afterthought reconstructed from unstructured message text.

Caching Helps, but Membership Data Can Become a Bottleneck

Slack's account of its February 2022 incident describes an overloaded Vitess keyspace containing channel membership data. It is a reminder that the hottest data may be the information needed to open the application, rather than the message bodies themselves. Slack's incident analysis

In a conceptual client boot sequence, the application asks who the user is, which channels they can access and where they last read. Many requests can converge on these shared dependencies.

Caching reduces repeated work, but a cache reset can expose the underlying load all at once. A recovery plan should account for that cold-start traffic. Limiting concurrent refreshes and avoiding unnecessary full reloads can matter as much as optimising the query itself.

This also explains why simply adding more application servers sometimes worsens an incident: more workers may generate more requests against the same constrained datastore.

Retention and Deletion Are Data-Lifecycle Problems

Imagine our fictional organisation retains messages for a fixed period. Expiring a message from the primary table is insufficient if search results continue to display its text.

The retention workflow must account for derived indexes, cached responses and referenced files. Different storage systems may complete cleanup at different times, but the serving layer needs clear rules about what it can still return.

Backups introduce another distinction. A retained recovery copy does not have to be directly available to ordinary search. Conversely, restoring a backup must not silently reactivate content that was removed afterwards. A restoration procedure needs to reapply relevant deletion state.

These are design considerations for any workplace archive, not a statement of Slack's exact retention internals. The central idea is that deletion is a workflow across representations, with observable completion and defined recovery behaviour.

Search Quality Requires More Than Finding Matching Words

Two messages may contain the same phrase while having very different value. One might be the final decision; another might quote an abandoned proposal. A search experience needs relevance as well as retrieval.

In a teaching design, useful signals might include textual match quality, recency and the searcher's selected filters. The team should evaluate these against real tasks, such as finding the last approved release plan, rather than assuming the newest match is always best.

Ranking should not obscure the result's context. Showing the channel, author and date helps the reader judge whether a message answers the question. A snippet without context can confidently surface a statement that was later reversed.

For engineers, this separates two concerns: the index makes candidates discoverable, while ranking and presentation help people decide which candidate matters. Improving one does not automatically improve the other.

Testing the Boundaries of the System

A useful test suite for our project channel would include sending during an indexing outage, editing a message several times quickly and removing channel access while search results are cached.

It should also test reconnecting after missing live events. The client must fill the history gap without duplicating messages it already received. Stable message identifiers make reconciliation easier than comparing text or timestamps alone.

For a storage migration, compare actual results and behaviour across old and new paths. Check unusual cases such as empty threads, deleted parents and old messages. A migration can preserve the number of rows while still changing how permissions or ordering work.

Operationally, measure durable-write latency, history-read latency, indexing delay and authorisation failures separately. A single green dashboard light cannot describe all the promises a workplace messaging product makes.

Rebuilding an Index Is a Product Operation

Suppose our search service needs a new document format. Rebuilding the index from message history can take time, and messages will continue changing during that process.

A safe teaching approach is to build a separate index from a known starting point, apply subsequent changes and compare representative queries before switching readers. The old index remains available until the replacement meets its freshness and correctness checks.

The difficult part is the boundary between the historical copy and the live change stream. Missing events at that boundary creates invisible gaps; applying updates without versions can restore old content over new edits.

Operationally, reserve capacity for the rebuild so it does not overwhelm ordinary history reads. Search maintenance should not make the conversation itself unusable. This is why derived data is only safely rebuildable when the team has designed and exercised the rebuild path, rather than assuming that keeping the original messages is sufficient.

The Big Picture

Slack's published designs show a separation between durable message storage, real-time distribution and searchable history. Partitioning supports growth, indexes support retrieval, and permission checks keep those capabilities aligned with the user's access.

The lesson for smaller systems is to make each promise explicit. A message can be saved before it is delivered, delivered before it is searchable, and searchable only within the right audience.

Once those boundaries are clear, queues, caches and storage choices become easier to evaluate. They are tools for preserving a useful conversation over time, including when networks fail, data changes and people join or leave the discussion.