I thought synchronized music would be the hard part.
It wasn’t.
As Vibe Village evolved, SharePlay became responsible for far more than keeping several devices on the same song. We began using Apple’s Group Activities framework as the coordination foundation for session chat, interactive Trivia, scoring, participant roles, playback authority, Room Audio ownership, and recovery when devices leave and return.
Somewhere along the way, I realized I wasn’t simply building a SharePlay enabled music app anymore.
I was building a distributed social experience on top of SharePlay.
That realization has changed how I think about mobile development.
SharePlay is still a relatively young platform
Apple introduced SharePlay at WWDC21 in June 2021 with an ambitious idea: FaceTime could become a place where people didn’t simply communicate. They could actually do things together.
Behind SharePlay was the new Group Activities framework. Apple demonstrated experiences such as watching movies, listening to music, exercising together, and sharing other activities while FaceTime provided the human connection around them.
The public rollout took a little longer. SharePlay ultimately became available with iOS 15.1 and iPadOS 15.1 in October 2021.
The first wave included some enormous names. Apple Music, Apple TV+, Apple Fitness+, TikTok, Twitch, Paramount+, the NBA and other services adopted SharePlay experiences.
But Apple’s vision was broader than synchronized media.
From the beginning, Apple showed developers how GroupSessionMessenger could exchange custom application data among participants. One of its early WWDC demonstrations transformed a drawing application into a real time shared canvas.
That was an important signal.
SharePlay could synchronize a movie or a song, but Group Activities could also become infrastructure for entirely new categories of multiuser applications.
The framework continued evolving. Apple later made it possible to initiate SharePlay without first establishing a FaceTime call and expanded the ways activities could begin from an application’s own interface.
Yet SharePlay still occupies an interesting corner of the Apple development ecosystem.
Most applications don’t need to reason about several devices participating in the same live experience. Fewer make that shared experience central to the product itself.
That leaves considerable territory for developers willing to explore what Group Activities can become beyond its most familiar uses.
That is where Vibe Village has gone.
We aren’t using SharePlay solely to let two people press play on the same song.
We’ve been building an application level social protocol on top of it.
SharePlay became a foundation, not the finished experience
When I first began working with SharePlay, synchronized media was the obvious capability.
Vibe Village eventually needed participants to share much more than playback position.
Listening Sessions, for example, include session chat. People participating in the experience can communicate while they listen. Those messages belong to the social experience surrounding the music, not to playback itself.
Trivia added another dimension.
Trivia isn’t simply content displayed independently on every device. The group needs a consistent understanding of where the Trivia experience is in its lifecycle.
Questions need to appear at the appropriate moments. Participants submit answers independently. The session has to coordinate question state, responses, results and scoring as the music experience progresses.
We therefore extended the experience built on SharePlay to coordinate several distinct kinds of shared information:
They may travel over the same underlying framework, but they do not have the same semantics.
A chat message is an event. A Trivia answer contributes to evolving application state. A Room Audio assignment establishes responsibility. A recovery request asks another participant to describe state that already exists. Playback synchronization describes where the group should be in time.
Treating all of those as generic “SharePlay messages” would make the system increasingly difficult to reason about.
One of the lessons has therefore been to define the meaning, authority, lifecycle and recovery behavior of each kind of shared state explicitly.
A framework can give you the communication mechanism without giving you the distributed application protocol your product needs.
SharePlay coordinates an experience. It doesn’t own your application state.
Being in the same GroupSession does not mean every device knows the same thing at the same moment.
Vibe Village has two social music experiences that make this particularly important. Music Battles allow people to compete with songs and vote on the results. Listening Sessions allow a group to experience music together without requiring a winner.
Both depend on shared state that extends well beyond playback.
Who is participating? Who has authority over a decision? What song is canonical? Is playback running or paused? Which device should produce the room’s audio? Which Trivia question is active? What is the current score? Has a returning participant recovered the current state, or is it still operating from information that existed before it disconnected?
Those questions belong to Vibe Village.
SharePlay provides the infrastructure that allows the devices to coordinate. The application still has to determine what that coordination means.
A social music app becomes a distributed system surprisingly quickly
You don’t need a hundred servers to encounter distributed systems problems.
Two iPhones will do.
Once two devices independently observe events and exchange messages, familiar problems begin appearing: message ordering, duplicate delivery, stale state, authority, convergence, idempotency, identity, recovery and race conditions.
Imagine Device A advances to the next song while Device B is temporarily reconnecting. Device B returns. Should it initialize playback? Should it ask Device A what is happening? What happens if participant information arrives before playback state? What if playback state arrives before the identity of the participant responsible for Room Audio? What happens when several of those events arrive almost simultaneously?
None of these situations is particularly exotic. They are normal consequences of a multi device application.
The mistake is assuming there will always be one correct event order. There often isn’t.
The application needs to converge on the correct state regardless of which reasonable event order occurs.
Logical playback and physical audio are different things
One of the most consequential architectural lessons came from something that sounds simple: Which device should play the music?
Consider several people sitting together in a room. They are participating in the same Listening Session, but playing the same song through every iPhone would be undesirable.
Perhaps one participant has an iPhone connected to the room’s speakers. That device should produce the audio while everyone else continues participating in the shared experience.
That led to an important distinction inside Vibe Village. The shared playback state is one concern. The physical renderer is another.
I think of the latter as Room Audio.
A device can participate in the canonical session without being responsible for producing its sound. Likewise, the participant coordinating aspects of the experience does not necessarily need to own the speakers.
Once that abstraction existed, the system became easier to reason about.
What should be playing?
Where should it be heard?
Those are independent questions.
Authority should not automatically imply ownership
Vibe Village has a concept called the Family Circle Owner (FCO). Among other responsibilities, that participant can serve as an authority within the shared experience.
An early temptation in a system like this is to allow authority to imply ownership of other responsibilities. That creates trouble.
Suppose another participant has deliberately been selected as the Room Audio device. The FCO temporarily leaves and later rejoins.
The FCO remains important to the session, but its return should not mean: “I’m back, so I own the speakers again.”
Yet that is exactly the kind of bug a distributed application can produce when authority and rendering are insufficiently separated.
We encountered variations of it during development. A returning FCO could incorrectly become the Room Audio device alongside the participant who already owned it. In another iteration, rejoining could reset Room Audio ownership to the FCO and even disturb active playback.
Locally, the returning device lacked information. Globally, however, nothing had changed. Another participant still owned Room Audio.
The returning FCO needed to discover and respect that existing state.
Authority determines who may make certain decisions. Ownership describes the result of one of those decisions. They should not be confused.
Joining and rejoining are fundamentally different operations
This became one of the biggest lessons from the project.
A participant joining a new experience may need initialization. A participant returning to an existing experience needs reconciliation.
Initialization says: “There is no state yet. Establish it.”
Reconciliation says: “State already exists. Discover it and converge with it.”
Treating rejoin as initialization can be destructive.
During development, a returning device could behave as though it needed to establish a fresh playback baseline. At another point, the FCO could reclaim Room Audio even though a secondary device already owned it.
Both behaviors made sense from the perspective of the returning device. Both were wrong from the perspective of the group.
The returning device lacked information, but absence of local information did not mean absence of authoritative state.
Unknown is not the same thing as nonexistent.
Recovery deserves its own protocol
Once we accepted that principle, another conclusion followed.
A returning device shouldn’t reconstruct authoritative state from whichever messages happen to arrive first. It should be able to ask for it.
Conceptually, a rejoining participant needs to be able to ask: “What is the authoritative Room Audio assignment for the session I just returned to?”
Another participant possessing valid state can respond. The returning device validates that response and installs the recovered assignment locally.
It does not manufacture a new assignment. It does not take ownership. It does not reset playback.
That distinction is important because recovery is not mutation.
A recovery message describes what already exists. A handoff changes what exists.
Conflating the two can create spectacularly confusing bugs.
Identity becomes complicated faster than expected
“Which user is this?” sounds like an easy question.
In a SharePlay application, it often isn’t.
There can be a transport identity representing a participant in the current GroupSession. There may be an application identity representing that person in your own system. There may also be stable identity concepts that need to survive changes to the underlying transport identity.
Those identities answer different questions.
When a message arrives, I may need to know which SharePlay participant actually sent it. When presenting the interface, I want the person’s recognizable name. When someone disconnects and reconnects, I may need to determine whether a new transport identity corresponds to someone the application already knows.
Using the wrong identity for the wrong purpose can make a perfectly valid assignment appear to belong to nobody.
The lesson has been to make identity domains explicit rather than treating every identifier as interchangeable.
Sometimes “Resolving…” is the correct UI
One bug produced an unexpectedly useful UX lesson.
After an FCO rejoined a Listening Session, Vibe Village displayed Room Audio: None.
But that wasn’t necessarily true. The device simply didn’t know the answer yet.
Changing the interface to Room Audio: Resolving… was more accurate. The application knew authoritative state should exist but had not recovered enough information to present it confidently.
Then we encountered an even more interesting problem. It stayed on Resolving… forever.
That was frustrating, but the interface was no longer lying. It was accurately exposing unresolved distributed state.
A distributed application sometimes needs at least three presentation states: Yes. No. I don’t know yet.
Collapsing the third into either of the first two may make an interface look simpler while making the application misleading and considerably harder to debug.
The visible failure may be nowhere near the actual defect
The permanent Resolving… state became one of my favorite examples from this work.
Initially, it looked like a presentation problem. Then it appeared to be assignment reconciliation. Then we investigated authority hydration, recovery timing, and message receiver registration.
Eventually, physical device diagnostics exposed something far more specific.
The responder could not decode the Room Audio recovery request.
We had created custom recovery request and response types for the protocol. Those wire types had been declared private.
The sender could encode the recovery request, but the receiving side could not resolve the typed decoder through GroupSessionMessenger.
The visible symptom was a label that said Resolving…. The actual failure was several architectural layers beneath it.
The fix itself was tiny. Making the recovery wire types module visible allowed the protocol to complete.
The returning FCO could request the existing Room Audio assignment. The other device could respond. The FCO could install the recovered state without taking ownership or disturbing playback.
The interface finally replaced Resolving… with the actual participant responsible for Room Audio.
In distributed applications, the place where a failure becomes visible may be nowhere near the place where the failure occurred.
A label that won’t update can ultimately be a transport problem.
Observability needs to describe the protocol
That decoder problem would have been considerably harder to isolate if our logs simply said: “Recovery failed.”
Distributed systems need more context than that.
We gradually instrumented the lifecycle of the protocol itself. Was recovery requested? Was the request transmitted? Did another participant receive it? Was that participant eligible to respond? Was a response sent? Did the requester receive it? Was the response accepted? Was the assignment installed? Did participant identity subsequently resolve? Which controller instance handled each event? Which playback run and assignment revision were involved?
That changed the debugging process.
Instead of asking, “Why does the screen say Resolving?” we could ask, “Where did the protocol stop progressing?”
That is a much better engineering question.
Physical devices remain part of the test architecture
Automated testing has been invaluable on Vibe Village.
We have regression coverage around playback, synchronization, participant behavior, Room Audio, rejoining and other parts of the Listening Session lifecycle.
But some bugs survived all of it.
They appeared only when two physical devices joined an actual SharePlay session, one left, the other continued playing, the first returned, participant metadata hydrated asynchronously, MusicKit interacted with the renderer, and messages crossed the session boundary in real time.
One particularly stubborn Room Audio defect repeatedly passed focused tests and simulator builds while continuing to fail on actual devices.
The physical devices eventually gave us the evidence the automated environment could not.
A simulator can tell me that my state machine behaves correctly under the event sequence I modeled. Physical devices have an irritating habit of introducing event sequences I didn’t model.
Both are necessary.
I’ve increasingly come to treat multi device physical testing as part of architecture validation rather than something performed after development is supposedly complete.
The architecture I wish I had started with
If I were designing the system again with everything I’ve learned, I would think about it in layers:
The important part is the direction of dependency.
The presentation layer shouldn’t manufacture ownership because it cannot find an owner. The renderer shouldn’t invent canonical playback because its local player lacks state. A returning participant shouldn’t initialize a session because its local copy disappeared.
Each layer should reconcile from authoritative information above it.
And when that information isn’t available yet, the system should preserve uncertainty rather than fabricate certainty.
SharePlay gave us more room to innovate than I expected
What has surprised me most is how far the Group Activities model can be pushed.
Apple explicitly designed GroupSessionMessenger for custom application data. Its developer guidance has demonstrated collaborative experiences extending beyond synchronized media.
That invitation to experiment is important.
For Vibe Village, SharePlay now helps provide the coordination fabric beneath music playback, Room Audio, participant state, session chat, Trivia and scoring.
Those features don’t exist because SharePlay automatically provides them. They exist because SharePlay gives us enough infrastructure to build them.
There is a meaningful difference.
We are among a relatively specialized group of developers exploring SharePlay as something more than synchronized media. I don’t know how large that group is, and I wouldn’t pretend there is a reliable public count.
What I do know is that there is still considerable room for experimentation.
SharePlay is mature enough to provide remarkable infrastructure, yet specialized enough that developers willing to go deeper can still explore what a truly social application built around it might become.
A glimpse of where collective intelligence could take this
What happens when a shared application begins to understand the group, rather than simply coordinate its members?
Building Vibe Village has also intersected with another area of my research: Collective-State Inference (CSI).
CSI explores a broader question about groups: whether meaningful collective states can be inferred from the relationships, interactions and patterns among participants rather than reducing the group to a collection of individual attributes.
That research is not currently implemented as part of Vibe Village’s SharePlay architecture. But building these shared experiences has made the connection increasingly interesting.
Consider what a Listening Session already produces conceptually. People join and leave. They choose whether they are together or remote. Someone assumes responsibility for Room Audio. Participants talk. They answer Trivia questions. They react to songs. Authority moves through the experience. Patterns emerge over time.
Today, the application coordinates those interactions.
A future generation of social applications could potentially reason about the collective state emerging from them.
Is the group highly engaged or beginning to disengage? Is participation balanced or dominated by one person? Are people converging around particular music? Did a Trivia question energize the room? Is the group behaving differently tonight than it typically does?
Those are not simply questions about individual users. They are questions about the group as a system.
That distinction opens an intriguing direction for Vibe Village. SharePlay can help establish the shared experience. Application telemetry can describe interactions within it. A future reasoning layer informed by ideas such as CSI could potentially help interpret what those interactions mean collectively.
There is considerable research and engineering between that possibility and a production capability, and I don’t want to imply otherwise.
The next generation of social applications may not simply connect people or synchronize their devices. They may become capable of understanding something about the state of the group itself.
SharePlay changed how I think about mobile development
I started by trying to keep a few iPhones on the same song.
That led to questions about authority, identity, ownership, recovery and eventually collective state.
Earlier, I said that you don’t need a hundred servers to encounter distributed systems problems.
Two iPhones will do.
I’ve learned something else along the way.
You don’t need to begin with an elaborate collective intelligence system to start asking interesting questions about groups. Sometimes you can begin with a few people, a few phones, music they love or music they’re discovering for the first time, and an experience they are creating together.
We’ve learned how to connect people. We’re learning how to coordinate their experiences. The next frontier may be understanding something about the group that emerges.
The technology connects the devices. The interesting part is what emerges among the people.