Building a Messaging App from the Device Out: Local-first, E2EE and On-Device Search
Most system design answers for a chat app take it for granted that the server holds the message history. Hand that job to the user’s devices and keep the server as a relay, and the design has to answer a different set of questions: where identity and keys live, what the server can still see, how messages get ordered, how a new device gets old messages, and where search runs. I’ll go through them in that order, pointing to what shipping apps have published along the way.
The usual answer follows a familiar outline: how to design DynamoDB GSIs, Kafka or Redis PubSub, ZooKeeper for connection state, consistent hashing to find the right chat server, one device ID per device, sequence numbers for multi-device sync. The syllabus of a system design course on Wondering is laid out exactly like that.
Decide who holds the data
The term local-first comes from Ink & Switch’s 2019 essay “Local-first software: You own your data, in spite of the cloud”, by Martin Kleppmann and three co-authors, presented at Onward! 2019. It sets out seven ideals: fast, multi-device, offline, collaboration, longevity, privacy and user control.
Read them as requirements for a messaging app and each one turns into a design decision:
| Ideal | What it means for a messaging app |
|---|---|
| Fast (no network round trip) | Writing a message hits the local database first; the UI doesn’t wait for the server |
| Multi-device | Every device keeps a full copy of the history |
| Offline | Read and write offline, send when the connection is back |
| Collaboration | Several people post in a group at once and merging needs no central coordinator |
| Longevity | If the service shuts down, local history can still be read and exported |
| Privacy | End-to-end encrypted, so the server can’t read content |
| User control | An open data format that users can take with them |
Longevity and user control are the hard ones. Local history survives, but as long as identity (a phone number or an account) and message relay depend on a central server, you can’t send or receive once it’s gone, and the vendor still decides the data format. Full decentralisation is a much bigger project, so this post assumes a relay server stays and moves everything else it can onto the device.
Identity and devices
Once the device holds the data, the first question is identity. Each device generates its own identity key and the private half never leaves it. The server keeps a directory of which devices belong to which account, with their public keys. To send a message, the sender encrypts a separate copy for every device on both sides (client fan-out).
WhatsApp’s multi-device design, published in 2021, is the clearest published example. In the old setup the phone was the only source of truth, and the web and desktop apps mirrored it over a persistent connection; when the phone dropped off, so did they. The new design gives every device its own identity key, the sender encrypts N copies for N devices, and the server keeps only the account-to-device mapping.
That makes the key directory the most sensitive thing on the server. The server hands out public keys, and if it swaps in one of its own, end-to-end encryption no longer means anything. The traditional defence is for two people to scan a QR code or compare a security code in person, but Meta worked out that a 100-person group would need 4,950 pairwise checks. On 13 April 2023 WhatsApp announced key transparency: the server keeps an append-only Auditable Key Directory, every change goes into a public audit record, and a client can check for itself that a contact’s key is the one everyone else sees. The implementation, akd, is open-source Rust.
What the server can see
What stops the server reading messages is end-to-end encryption (E2EE). Every app claims to encrypt, but the guarantees vary a lot, so it’s worth separating a few terms that tend to get blurred.
Transport encryption versus end-to-end encryption
Transport encryption (TLS and the like) protects the hop from the phone to the server; the server decrypts what arrives, stores it, and can read it. With end-to-end encryption the keys exist only on the participants’ devices and the server relays ciphertext it can’t open. Telegram’s ordinary cloud chats are the first kind: messages are stored on the server without E2EE by default. Only one-to-one secret chats are end-to-end encrypted, and each side’s copy is tied to the device that took part, so switching devices loses them. Defending the cloud chat design, Durov argued that server storage avoids insecure third-party backups and lets users reach their messages from any device.
Forward secrecy
A key leaked today doesn’t unlock yesterday’s messages. The usual approach is short-lived session keys that are thrown away after use, so the current key can’t be used to work out earlier ones. The Double Ratchet goes further and uses a new key for every message.
Post-compromise security
This is the guarantee in the other direction: after a device is compromised and its keys leak, the conversation becomes secure again once the two sides complete a fresh key exchange. Some call it future secrecy. The Double Ratchet specification gives the two properties to different parts. The symmetric-key ratchet makes past keys look random to an attacker, but its inputs are fixed, so it can’t recover from a break-in. The Diffie-Hellman ratchet brings in a new key pair each round and supplies the recovery.
Metadata
E2EE protects content, not who talked to whom and when. Wikipedia’s entry on the Signal Protocol states plainly that it doesn’t provide anonymity. The account-to-device directory from the previous section is metadata the server will always have.
Post-quantum
The threat here is “record the ciphertext now, decrypt it once quantum computers are good enough”. Where the main players stand:
- Signal introduced PQXDH in 2023, replacing the initial key agreement with a post-quantum version.
- Apple announced iMessage PQ3 on 21 February 2024, shipping in iOS 17.4. Besides the initial key agreement, it renegotiates with Kyber-768 during a conversation, roughly every 50 messages and at least every seven days. Apple sorts messaging apps into Levels 0 to 3 and puts PQ3 at Level 3, which is Apple’s own grading.
- Signal announced the Triple Ratchet on 2 October 2025: a Sparse Post-Quantum Ratchet (SPQR) using ML-KEM 768 runs alongside the existing Double Ratchet, and the outputs of both are mixed into the message keys. An ML-KEM encapsulation key is 1,184 bytes and a ciphertext 1,088 bytes, against 32 bytes for ECDH, so SPQR uses erasure coding to cut the large keys into small pieces carried across several messages.
Groups
Groups cost far more than one-to-one chats. The simplest approach is pairwise fan-out, a separate copy for each member; the client fan-out above works the same way, and the cost grows linearly with members and devices. The alternative is sender keys: each sender has one group key and encrypts once for everyone. RFC 9420 notes that this gives forward secrecy, but achieving post-compromise security makes the update cost grow with the square of the group size.
Matrix’s Megolm is a sender-key scheme. The sender creates an outbound session and hands its key to every device in the room over one-to-one Olm channels; after each message the key is hashed forward, so anyone holding the current state can read later messages but not earlier ones. The IETF published Messaging Layer Security (MLS, RFC 9420) in July 2023. It encapsulates keys in a tree, so the cost of deriving and updating the shared key grows with the logarithm of group size, and the RFC describes it as suitable for groups “ranging from two to thousands”.
What existing apps chose
Here is what the apps themselves have published:
| App | Protocol | E2EE by default | Notes |
|---|---|---|---|
| Signal | Signal Protocol (PQXDH, Triple Ratchet) | Yes | SPQR is rolling out gradually; making it mandatory needs a later update |
| Signal Protocol | Yes, since 5 April 2016 | Key transparency added in 2023; E2EE for cloud backups is opt-in (since 14 October 2021) | |
| Messenger | Signal Protocol plus Labyrinth | Personal messages and calls, rolled out from 6 December 2023 | Message history stored encrypted on Meta’s servers |
| iMessage | PQ3 | Yes | PQ3 since iOS 17.4 |
| RCS (Google Messages, iPhone) | Signal Protocol for one-to-one chats in Google Messages (since November 2020); MLS across platforms under GSMA Universal Profile 3.0 | On by default on iPhone from iOS 26.5, in beta from 11 May 2026 | Needs carrier support; Apple’s announcement doesn’t name the protocol, the MLS detail comes from GSMA’s description of Universal Profile 3.0 |
| LINE | Letter Sealing v2 (ECDH, AES-256-GCM) | Yes, on by default in the main clients since 2016, can’t be turned off since 2021 | One-to-one chats, groups of up to 50, one-to-one calls; Open Chat, official accounts and group calls aren’t covered, and a chat loses protection once a bot joins |
| Telegram | MTProto 2.0 | No | Only one-to-one secret chats are E2EE; groups never are |
| Matrix (Element and others) | Olm/Megolm (vodozemac) | Private conversations, since May 2020 | MLS for group encryption planned under MSC2883 |
Instagram, missing from the table, went the other way. In March 2026 Meta announced that end-to-end encryption for Instagram direct messages would stop on 8 May 2026, saying only a small share of users had switched it on. On Instagram it had always been opt-in, and users who wanted encrypted chats were pointed to WhatsApp, where it’s the default. Within the same company, WhatsApp and Messenger turned it on for everyone and Instagram never did.
LINE matters more here because of its reach in Taiwan. According to the LINE-Break talk at Black Hat, it has about 21 million users in Taiwan, roughly 90% of the population. In 2017 researchers from the Citizen Lab at the University of Toronto and the University of New Mexico presented an analysis at FOCI, looking at version 6.7.1. They found two problems: the MAC didn’t cover metadata such as sender, recipient and sequence number, so recorded messages could be replayed; and forward secrecy only applied between client and server, not between clients. LINE accepted both points and explained the second by the need to sync with secondary devices. LINE’s 2021 encryption whitepaper, v2.1, still places forward secrecy only in the transport layer between client and server.
In 2025 Diego F. Aranha and colleagues at Aarhus University in Denmark published LINE-Break (Black Hat Europe 2025, with the paper in ASIACCS 2026), this time on Letter Sealing v2. They concluded that there is no continuous key rotation, so no forward secrecy by design; that messages can be replayed, reordered or dropped without either end noticing; that a malicious user colluding with an attacker can forge who sent a message; and that stickers and URL previews leak more plaintext than the official documentation says. The researchers reported the findings on 6 June 2025 and LY confirmed them. LY’s response said it had chosen to order messages by the time the server received them, for the sake of user experience, and that it was looking into updating the protocol.
Ordering: who issues the sequence number?
In the usual answer, the server assigns an increasing sequence number as each message arrives. Every device then sees the same global order, but without the server there’s no order at all, and anything written offline has to wait. LINE’s server-timestamp ordering, mentioned above, is this approach.
The local-first approach lets each device write on its own and merges later. If the merge function is commutative, associative and idempotent, the final state is the same however many times and in whatever order updates arrive. Those are the conditions for a state-based CRDT, formally defined by Shapiro and colleagues in 2011. A chat log is about the easiest structure to fit: assume messages are only ever added, never edited or deleted, and merging is a set union followed by a sort on a deterministic key.
The code below uses a Lamport clock plus device ID as that key. Two devices each write two messages offline, then merge with each other. The test messages are in Chinese: the phone has “bring a raincoat for Saturday’s hike” and “the trailhead meeting time is now 7am”, the laptop has “the tax deadline is the end of May” and “the flight is the early one out of Taoyuan”.
from dataclasses import dataclass
from importlib.metadata import version
import platform
import tempfile
from fastembed import TextEmbedding
from qdrant_client import QdrantClient, models
@dataclass(slots=True, frozen=True)
class Message:
lamport: int
device: str
text: str
@property
def key(self) -> tuple[int, str]:
return (self.lamport, self.device)
class DeviceLog:
def __init__(self, device: str) -> None:
self.device = device
self.clock = 0
self.messages: dict[tuple[int, str], Message] = {}
def write(self, text: str) -> None:
self.clock += 1
msg = Message(self.clock, self.device, text)
self.messages[msg.key] = msg
def merge(self, other: "DeviceLog") -> None:
# Union of logs: commutative, associative, idempotent.
self.messages.update(other.messages)
self.clock = max(self.clock, other.clock)
def ordered(self) -> list[Message]:
return sorted(self.messages.values(), key=lambda m: m.key)
phone, laptop = DeviceLog("phone"), DeviceLog("laptop")
phone.write("週六爬山,記得帶雨衣")
phone.write("登山口集合時間改成早上七點")
laptop.write("報稅截止日是五月底")
laptop.write("這次機票訂桃園出發的早班機")
phone.merge(laptop)
laptop.merge(phone)
assert phone.ordered() == laptop.ordered()
for m in phone.ordered():
print(m.lamport, m.device, m.text)
Output:
1 laptop 報稅截止日是五月底
1 phone 週六爬山,記得帶雨衣
2 laptop 這次機票訂桃園出發的早班機
2 phone 登山口集合時間改成早上七點
Both devices end up with the same order and the assertion passes. But the laptop’s and the phone’s messages are interleaved in a way that has nothing to do with when they were written. A Lamport clock gives a total order consistent with causality, not wall-clock time, so in a chat, things two people said while offline get shuffled together after the merge. Having the server issue sequence numbers avoids that oddity, at the price of no ordering while offline. Whether ordering belongs on the server or the device depends on how much the product cares about writing offline.
History on a new device
Once encryption and ordering are settled, the next question is where old messages come from when a new device joins. Existing apps have made quite different choices:
- Telegram’s cloud chats keep history on the server in a form the server can decrypt, so any device can log in and see everything.
- LINE gave up client-to-client forward secrecy, citing sync with secondary devices (see above).
- WhatsApp has the primary device encrypt a bundle of recent chats and send it straight to the new device, so the data stays among the user’s own devices.
- Messenger uses Labyrinth to store encrypted history on Meta’s servers, with keys under the user’s control. One of the problems Meta listed was letting people read their history without relying on local storage.
Forward secrecy wants old keys deleted; history wants a new device to read old messages later. One of them has to give. Either old messages are re-encrypted somewhere under a long-term key, or they can only come from another device that still has them in plaintext. WhatsApp’s device linking and local-first take the second route and treat the devices as the source of truth. Messenger takes the first, and its servers hold only ciphertext they can’t read. WhatsApp’s cloud backup also falls into the first group, except that it isn’t end-to-end encrypted unless the user turns that on.
A copy on someone’s server also carries legal exposure. In February 2025 the UK government reportedly used the Investigatory Powers Act to demand that Apple provide access to encrypted iCloud data. Apple then withdrew Advanced Data Protection for UK users, existing users were given a deadline to turn it off, and iCloud backups and similar data went back to standard protection, where Apple holds the keys. According to The Register’s report of 24 February 2025, iMessage’s own end-to-end encryption was unaffected.
In local-first terms, E2EE covers the privacy ideal, and the side effect is that the server can do less for you. If it can’t decrypt, it can’t search, merge or index on your behalf, and all of that work moves to the device.
Search has to run on the device
A local database handles keyword search easily. Semantic search (“what time did we say to meet?”) needs a vector index on the device. Qdrant offers two routes:
- Local mode in
qdrant-client:QdrantClient(path=...)runs in the same process and stores data in the folder you name, with no separate service. The README lists development, prototyping and testing as its uses. - Qdrant Edge: announced on 29 July 2025 as embedded vector search for robots, phones, point-of-sale terminals and IoT. It runs as a library with no background optimiser or update threads, so every operation is synchronous and under the application’s control. It launched as a private beta; the official docs now mark it as beta and warn that the API and features may change. The Python package is
qdrant-edge-pyand there’s also a Rust crate,qdrant-edge.
Continuing the code from the ordering section, this turns the merged messages into vectors with a multilingual embedding model and searches them in local mode. The query is “上次說幾點集合?”, roughly “what time did we say to meet last time?”:
# Pin the model name; every device must embed with the same model and version.
MODEL = "sentence-transformers/paraphrase-multilingual-mpnet-base-v2"
embedder = TextEmbedding(MODEL)
docs = phone.ordered()
vectors = list(embedder.embed([m.text for m in docs]))
with tempfile.TemporaryDirectory() as path:
client = QdrantClient(path=path) # in-process, data stays in this folder
client.create_collection(
"chat",
vectors_config=models.VectorParams(
size=len(vectors[0]), distance=models.Distance.COSINE
),
)
client.upsert(
"chat",
points=[
models.PointStruct(
id=i,
vector=v.tolist(),
payload={"text": m.text, "device": m.device, "lamport": m.lamport},
)
for i, (m, v) in enumerate(zip(docs, vectors))
],
)
query = next(iter(embedder.embed(["上次說幾點集合?"]))).tolist()
hits = client.query_points("chat", query=query, limit=2).points
for h in hits:
print(f"{h.score:.3f}", h.payload["text"])
client.close()
print("Python", platform.python_version())
print("qdrant-client", version("qdrant-client"), "fastembed", version("fastembed"))
Output:
0.538 登山口集合時間改成早上七點
0.390 週六爬山,記得帶雨衣
Python 3.13.12
qdrant-client 1.19.1 fastembed 0.8.1
The top hit is the trailhead meeting time, followed by the raincoat reminder. The first attempt used the smaller paraphrase-multilingual-MiniLM-L12-v2 (about 0.22 GB according to fastembed). For the same question it ranked the tax deadline first (0.404) and the right answer second (0.322); only the roughly 1 GB paraphrase-multilingual-mpnet-base-v2 above got it right. Four messages prove nothing about either model. Evaluate on your own data, and on a phone, expect to trade file size against accuracy.
When each device builds its own index, every device has to produce the same embeddings. Pin the model name and version, the tokeniser and the normalisation, or a message may turn up in search on one device and not on another. The same goes for hand-rolled features: hash with hashlib, because Python’s built-in hash() is salted per process and gives different results on the phone and the laptop for the same text.
The other decision is whether to sync the index. A vector index is derived data that can be rebuilt from the messages, so the simpler option is to sync only the message log and let each device build its own index. The cost is a full rebuild when a device is linked, and you’ll need to measure what that does to battery and time on a phone.
What the server still does
The server’s job shrinks to a few things:
- Device directory: which devices belong to each account and their public keys, plus the key transparency audit record clients use to check it.
- Mailbox: holding encrypted messages while the recipient’s device is offline, deleting them once delivered.
- Push notifications to wake devices up.
The DynamoDB GSIs, ZooKeeper and consistent hashing from the usual answer are still useful, but they now serve encrypted envelopes waiting for delivery rather than chat history. There’s far less data and it’s kept for far less time, so the design effort shifts from query patterns to connection handling and delivery guarantees.
The complexity the server sheds lands on the client: merge logic, local schema migrations, on-device indexing, and recovery when a device is lost.
Notes
- The usual messaging design assumes the server owns the history. Move that to the device and the server is left with a device directory, an offline mailbox and push.
- Every device has its own identity key and the server holds the key directory, which is why key transparency matters: clients can check the directory hasn’t been tampered with.
- Transport encryption, E2EE, forward secrecy and post-compromise security are different guarantees. LINE-Break found that Letter Sealing v2 has no forward secrecy by design.
- Group encryption costs differ widely: pairwise is linear, sender keys go quadratic once you want post-compromise security, MLS is logarithmic.
- A Lamport clock plus device ID gives every device the same merged order, but not wall-clock order.
- Data the server can’t decrypt can only be searched on the device. Pin the embedding model and evaluate it on your own data.
Further reading
- Wondering, How to Design WhatsApp: a system design course; only its public syllabus is cited here, as an example of the usual answer.
- Ink & Switch, Local-first software: You own your data, in spite of the cloud: the 2019 essay and the source of the seven ideals.
- Meta Engineering, How WhatsApp enables multi-device capability: the 2021 first-hand account of the multi-device architecture.
- Meta Engineering, Deploying key transparency at WhatsApp: 13 April 2023, the AKD design and the 4,950-check example.
- Meta Engineering, Building end-to-end security for Messenger: 6 December 2023, Labyrinth and multi-device history; the default rollout was announced in Launching Default End-to-End Encryption on Messenger.
- Help Net Security, Meta ditches end-to-end encrypted messaging on Instagram: 16 March 2026, the end date and Meta’s reasons, as reported in the press.
- Signal, The Double Ratchet Algorithm: the official specification and the source for how forward secrecy and break-in recovery are split.
- Signal, Signal Protocol and Post-Quantum Ratchets: 2 October 2025, the first-hand account of SPQR and the Triple Ratchet.
- Apple Security Research, iMessage with PQ3: 21 February 2024; the Level grading is Apple’s own.
- IETF, RFC 9420: The Messaging Layer Security (MLS) Protocol: the standard; the cost comparison of the three group approaches is in its Introduction.
- Matrix.org, End-to-End Encryption implementation guide: official docs on Megolm outbound sessions and key sharing.
- Privacy Guides, Apple Introduces End-to-End Encrypted RCS Messaging in the iOS 26.4 Beta: 19 February 2026, a report from the first iOS 26.4 beta citing GSMA on MLS in Universal Profile 3.0; secondary.
- Apple, End-to-end encrypted RCS messaging begins rolling out today in beta: the press release of 11 May 2026, which doesn’t name the protocol.
- Espinoza et al., Analysis of End-to-End Encryption in the LINE Messaging Application: FOCI 2017, peer-reviewed, covering the version of the time.
- Aranha, Hansen and Mogensen, LINE-Break: Cryptanalysis and Reverse Engineering of Letter Sealing: Aarhus University, ASIACCS 2026; the Taiwan user figure, the forward secrecy finding and LY’s response are in the Black Hat Europe 2025 slides.
- LINE, LINE Encryption Overview Technical Whitepaper v2.1: the official whitepaper of November 2021, vendor-written; the latest report is the LINE Encryption Report (2025).
- The Register, Apple ends iCloud Advanced Data Protection for UK customers: 24 February 2025, news report on the UK withdrawal.
- Qdrant, Qdrant Edge: Vector Search for Embedded AI: the announcement of 29 July 2025; current status in the Qdrant Edge docs. Both vendor-written.
- Wikipedia, Signal Protocol, Telegram (software), Matrix (protocol) and Conflict-free replicated data type: protocol components, timelines, MSC2883 and the CRDT definition; secondary, with primary sources in each entry’s references.
Ideas and technical judgement by Sheng; drafted with Claude · examples run on Python 3.13.12, qdrant-client 1.19.1 and fastembed 0.8.1