What it took to move a collaborative browser IDE beyond process memory
The first collaboration model in CodeVerse was convincing in exactly the way a local demo needs to be convincing. Open two tabs. Join the same room. Type in one editor. Watch the other editor update. Then ask one unpleasant question: what happens when those two sockets land on different server instances? The answer was that the room stopped being a room. Each process had its own memory, its own…
CodeVerse began its collaboration journey with a local demo that showcased the ability to open two tabs, join the same room, and type in one editor while the other editor updated in real-time. However, the demo quickly revealed a critical flaw - when those two sockets landed on different server instances, the room ceased to be a cohesive entity. Each process had its own memory, presence list, and idea of current files, leading to issues with state management during restarts, reconnects, and load balancing.
The real challenge was not just about Socket.IO's connection handling and room fan-out approach, but determining where the truth lived within the collaboration process. CodeVerse needed to address four key aspects: document state, room policy, presence, and durability. To achieve this, the team turned to Yjs for convergent document updates, Redis for live distributed room state and pub/sub, and Supabase for durable room snapshots and membership data.
Redis played a crucial role in CodeVerse's architecture, serving three distinct purposes. First, the Socket.IO Redis adapter enabled cross-instance fan-out, allowing updates from instance A to reach sockets connected to instance B without the need for a custom relay protocol. Second, Redis stored the normalized room record, containing files, active file, edit policy, organizer, encoded Yjs state, revision, and update time.
This live state was managed using expiring per-socket records in a sorted-set index, ensuring abandoned sockets did not remain visible indefinitely. Finally, Redis facilitated coordination by serializing compound room mutations with a small Redis lock acquired through SET NX PX. This lock prevented accidental removal of room state during expired lock scenarios, ensuring a safe and bounded process.
While Yjs excelled in convergence, it did not handle authorization. Every collaboration update had to pass server-side checks before reaching the Yjs document. The socket must still belong to the requested room, and a viewer could not edit, while a collaborator could edit only if permitted by the room policy. The organizer held the power to change the policy or remove collaborators, but clients could not grant themselves the organizer role by changing their payload.
After authorization, the server decoded the Yjs update, applied it to the stored document, and stored the new encoded CRDT state, incrementing the room revision. This separation of convergence and permission ensured that unauthorized edits remained unauthorized.
Network retries, browser reconnections, and user actions like double-clicking led to the need for operation IDs in each CRDT update. Redis recorded these operation IDs with an NX and an expiry, allowing only the first server instance to process the operation and mark it as successful. Subsequent deliveries would receive a successful duplicate acknowledgement without applying the edit twice.
This approach provided a practical idempotency boundary, ensuring at-least-once transport behavior with one accepted mutation per operation ID during the configured window.
Reconnect behavior posed a significant challenge as a socket ID was not a user identity and could change when the connection changed. CodeVerse addressed this issue by issuing an expiring signed reconnect token bound to the room, user ID, and role. When a room join succeeded, CodeVerse returned the latest CRDT snapshot and revision, allowing the client to recover the identity and session.
This approach made reconnect behavior testable, ensuring that a sample of clients reconnecting, potentially to the other application instance, would be acknowledged with their recovered identity. Durable snapshots were used sparingly, debounced to minimize latency and cost, while live updates relied on the more efficient Redis-based system.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.