Spatial Computing Interfaces: Design Rules and the ASO Playbook
Meta ships 69% of the eyewear, Apple ships the design language, Google ships the operating system. The interface rules are converging — and the app store rules for spatial software are not what mobile taught you.

Spatial computing interfaces are user interfaces that live in the room instead of on a screen: windows anchored to walls, objects that sit on tables, controls that respond to where you look, what your hands do and what you say. For two decades the design question was how to fit software into a rectangle. The question now is how software should behave when the rectangle is gone — and, for anyone building it, how a spatial app gets found in a store that has no bestseller list. The hardware market answered part of that in 2026, and not in the direction the headset makers expected.
AI Overview
A spatial computing interface places digital content in three-dimensional physical space and takes input through eye tracking, hand gestures, voice and accessories rather than a mouse, keyboard or touchscreen. In 2026 three platforms set the rules: Apple's visionOS, which established gaze-and-pinch as the default selection model; Google's Android XR, launched on Samsung's Galaxy XR with Gemini as a voice-first layer; and Meta's Horizon OS, which holds the largest installed base. IDC's June 2026 data shows the market tilting toward glasses — 13.6 million display-less and 3 million display glasses projected for 2026 against 3.2 million mixed reality headsets — which pushes interface design toward voice and audio. Distribution is the overlooked layer: the visionOS App Store has no curated top charts, so app store optimization for spatial apps is a search-metadata problem first and a screenshot problem second.
Key Facts
| Category | Technology — Design and Distribution Guide |
| Platforms | Apple visionOS 26 (visionOS 27 in developer beta), Google Android XR (Samsung Galaxy XR, $1,800), Meta Horizon OS |
| Input models | Gaze + pinch (visionOS default); voice via Gemini (Android XR); hand tracking and controllers (Meta); PS VR2 Sense controllers and Logitech Muse (visionOS 26) |
| 2026 shipments (IDC) | 13.6M display-less glasses · 3.0M display glasses · 3.2M mixed reality headsets |
| Market share, Q1 2026 | Meta 69.2%, RayNeo 3.4%, Xiaomi 3.1%, Viture 2.5%, XREAL 2.0% |
| Store constraints (visionOS) | 30-char name, 30-char subtitle, 3840×2160 screenshots, ~30 s preview, no curated top charts |
| Original framework | The Spatial Interface Stack — five layers, five owners |
| Updated | September 9, 2026 |
Why It Matters
Interface transitions reset platform economics. The mouse made Microsoft; the touchscreen made Apple's services business and killed a decade of handset incumbents. Spatial input is the next reset, and the 2026 data says the winning form factor is not the one the design language was written for. Apple wrote the grammar — gaze, pinch, floating glass windows — on a $3,499 headset; Meta ships 69% of the eyewear at a $376 average selling price, most of it with no display at all. The developers who win the transition are the ones whose interface survives being ported from a headset with eye tracking to a pair of glasses with a microphone and a speaker. For operators this is a decision about which layer of the stack to own; for investors it is a reminder that the installed base and the design leadership are, for now, at different companies.
Who Sets the Interface Rules in 2026?
Three platforms, three answers to the same question.
Apple — visionOS. Apple's Human Interface Guidelines for visionOS define the vocabulary most spatial designers now use: windows that float at a comfortable distance, "glass" materials that let the room show through, targets sized for gaze precision, and selection by looking then pinching. visionOS 26 added the pieces a static design language lacked: widgets that persist in the room and reappear each time the device is worn, spatial scenes generated from ordinary photos, shared experiences for users in the same room with remote participants joining over FaceTime, a Protected Content API for enterprises, and support for PlayStation VR2 Sense controllers and the Logitech Muse stylus. The visionOS 27 developer beta pushes realism further — virtual objects that cast light on real surfaces, real-time cloth simulation, a Reverb Mesh API that models how sound scatters off room materials, and a Foveated Streaming framework that lets cloud-rendered content arrive at full quality only where the user is looking.
Google — Android XR. Android XR shipped on the Samsung Galaxy XR at $1,800, with dual 3,552×3,840 micro-OLED panels, eye and hand tracking, and optional $250 controllers. The interface bet is different from Apple's: Gemini is the primary layer. Users navigate Maps in 3D by asking, get context about what is on screen by asking, and have flat photos and videos "auto-spatialized" into 3D on the fly. Existing Android apps run as flat panels; native XR apps get depth. Google is treating spatial as an extension of the phone ecosystem it already owns, and voice as the input that scales down to glasses.
Meta — Horizon OS. Meta does not lead on interface vocabulary; it leads on volume. IDC's Q1 2026 data puts Meta at 69.2% of the eyewear market, with the Ray-Ban line driving 2.25 million display-less glasses in the quarter — 167% year-over-year growth. On those devices there is no window to place and no gaze to track; the interface is a voice assistant, a camera and audio. That constraint, not visionOS's glass panels, is what most spatial software will actually run on by 2030.
Why Are Glasses Beating Headsets — and What Does It Do to the Interface?
IDC's December 2025 release marked the turn: XR shipments up 41.6% for 2025 while mixed reality headsets fell 42.8%, with the rest of the category up 211%. The June 2026 update projects the split holding — 13.6 million display-less glasses and 3 million display glasses in 2026 against 3.2 million headsets — and forecasts headsets recovering to 10.4 million by 2030 while display glasses grow faster, at a 41.9% compound rate, to 12.2 million. Average selling prices are forecast to fall 40% by 2030, to $229.
The interface consequence is direct. On a headset the designer controls the whole visual field and can assume eye tracking; on display glasses the field of view is narrow and content must be brief and peripheral; on display-less glasses the interface is entirely audio and voice. An app that hard-codes "look at the button, pinch" cannot ship to more than 80% of the devices IDC expects in 2026. An app whose core loop is "ask, hear, confirm" can ship to all of them, and gain a visual layer where the hardware offers one. Jitesh Ubrani's framing in the IDC note — that the real competition is "platform, ecosystem, and AI," not hardware — is the same conclusion from the vendor side.
This is where spatial computing crosses into the on-device AI economics we covered in Apple Intelligence and on-device AI: a voice-first interface on glasses is only as good as the model that runs within the thermal and battery budget of a frame light enough to wear all day.
The Spatial Interface Stack
Most spatial design advice is a list of tips — keep windows in the comfort zone, respect occlusion, combine pinch with voice. The tips are correct but they hide the structure. Spatial software has five layers, each owned by a different party, and most bad spatial apps are one layer making assumptions that belong to another.
| Layer | What it decides | Who owns it | Failure when ignored |
|---|---|---|---|
| 1 · Sensing | What the device knows about the room, eyes and hands | Hardware vendor | App assumes eye tracking that glasses do not have |
| 2 · Input grammar | How intent is expressed: gaze+pinch, voice, controller | OS | App invents its own gestures; users never learn them |
| 3 · Placement | Where content sits, at what depth, how it respects walls and furniture | App developer | Windows at fatiguing angles; objects that clip through tables |
| 4 · Persistence | Whether content survives removing the device or moving rooms | OS + app | State resets every session; widgets that forget their anchor |
| 5 · Distribution | How the app is found, previewed and installed | Store | Metadata written for a browse-driven store that does not browse |
Three rules follow from the stack.
Design at the input layer you cannot control. Layer 2 belongs to the OS. visionOS users already know gaze-and-pinch; Android XR users are learning to ask Gemini; Meta users know the controller. An app that adopts the platform grammar is learnable in seconds. An app that ships a custom gesture vocabulary is asking every user to attend a training session, and the store review will say so.
Placement is ergonomics, not aesthetics. Apple's guidelines put primary content slightly below eye level at roughly arm's length for a reason: neck and eye fatigue are the churn mechanism of spatial software. Occlusion — content disappearing behind a real wall — is not a visual flourish; it is the cue that tells the vestibular system the object is real, and skipping it is what makes long sessions nauseating.
Persistence is the product. visionOS 26's spatial widgets matter because they are the first mainstream example of layer 4 done by the OS: content that is still on the wall tomorrow. Before persistence, a spatial app is something you launch. After it, a spatial app is something that lives in your house — and that is the retention mechanism every other layer serves.
How Does ASO Work When the Store Has No Charts?
Layer 5 is where mobile intuition fails. On the iOS App Store, browse — top charts, editorial, category pages — drives a large share of installs. The visionOS App Store, as SplitMetrics' guide documents and as anyone opening it can confirm, lacks curated top charts, so users arrive through search and through Apple's own features. That changes the weighting of every metadata field.
Name and subtitle carry the keyword. Both fields are 30 characters. With no chart to be discovered on, the primary spatial keyword — the thing a user would say to search — belongs in the name or subtitle, not buried in the description. "Immersive," "spatial" and "3D" are the terms that distinguish native spatial apps from compatible iPad apps in search results.
Screenshots must be spatial, and sized for it. Apple's screenshot specifications require 3840×2160 pixels for Apple Vision Pro. The productive convention is to show the app in a real room — a kitchen counter, a desk — so a searcher understands placement before installing. A rendered void communicates nothing about layer 3.
Lead the preview with the interaction. Keep the video near 30 seconds and open with a hand pinching, a voice command being answered, or an object being placed. The user is evaluating the input grammar, not the logo.
Write for voice search. Inside a headset, typing is expensive and dictation is free. Metadata should contain the natural-language phrases people say — "watch movies in 3D," "measure my room" — rather than the compressed keyword strings mobile ASO trained everyone to write.
Ship 3D content the store can render. Apps built with RealityKit and USDZ assets can surface real objects in previews and in Safari's spatial browsing mode, which visionOS 26 opened to web developers for embedding 3D models. A store listing that links to a page showing the object in the user's own room is a preview no screenshot matches.
Protect the performance signals. Store ranking systems weigh crash rates and session quality on every Apple platform, and a spatial app that drops frames produces a physical symptom — discomfort — that mobile apps never could. Frame-rate stability is an ASO variable, not only an engineering one.
On Android XR the store is Google Play, where existing Android listings surface as flat panels; the same logic applies in reverse — a native XR listing needs to say so in its title to be distinguished from the flat version of itself.
Where Does Spatial Computing Meet Robotics and AI?
The sensing layer is shared infrastructure. The same depth cameras, hand tracking and scene understanding that let a headset place a window on a wall let a robot understand a kitchen; we made that argument in sharing space with machines, and visionOS 27's object tracking — metric-space pose, 90 Hz accessory tracking, extended training for hand-held objects — is essentially a robotics perception stack running on a consumer device. The generative layer is also converging: Google's auto-spatialization of flat media and Apple's generative spatial scenes both infer depth with a model, which means the content library for spatial devices is every photo and video ever taken, not the small catalog shot on spatial cameras. That is the network effect that has been missing, and it arrives through generative AI rather than through hardware.
For the platform contest itself, the relevant lens is the one we applied in platform economics: the real moat — the winner is whoever gets developers to design for their input grammar first, because grammar, once learned, is the switching cost.
Limitations
Shipment and share figures are IDC estimates and forecasts, which the firm revises quarterly; the 2030 projections in particular assume price declines and product launches that have not happened. Platform feature descriptions come from vendor documentation and launch coverage, not from independent testing, and visionOS 27 features are from a developer beta that may change before release. The App Store guidance reflects metadata rules and store behavior as documented by Apple and SplitMetrics at time of writing; Apple can add browse surfaces to the visionOS store at any point, which would shift weight back toward ratings and editorial. The Spatial Interface Stack is an analytical framework from the author's product work, not a survey-derived taxonomy, and the layer boundaries — especially between persistence and placement — are judgment calls.
The Bottom Line
Spatial computing interfaces have a settled grammar and an unsettled body. Apple wrote the rules for gaze and pinch on a device most people will never buy; Meta is shipping the devices most people will buy with no gaze to track; Google is betting that voice, through Gemini, is the layer that spans both. The design answer is to build at the input layer you cannot control — adopt the platform's grammar, design placement for fatigue, treat persistence as the product. The distribution answer is to stop thinking like a mobile marketer: in a store with no charts, the name field is the chart. Get the stack right and the same app runs on a $3,499 headset and a $229 pair of glasses. Get it wrong and it runs on neither for long.
References
- Smart Glasses Surge: The XR Market Is Rewriting Its Own Rules — IDC, June 15, 2026
- Global XR Shipments Rebound Behind Glasses-First Momentum — IDC, December 11, 2025
- visionOS 26 introduces powerful new spatial experiences for Apple Vision Pro — Apple Newsroom, June 2025
- What's New — visionOS — Apple Developer (visionOS 27 beta)
- Designing for visionOS — Apple Human Interface Guidelines
- Screenshot specifications — App Store Connect Help
- Adding 3D content to your app — Apple Developer Documentation
- Develop with Android XR — Android Developers
- Samsung Galaxy XR: price, specs and release date — Road to VR
- The Ultimate App Store Optimization Guide for the visionOS App Store — SplitMetrics
Related
Part of our ongoing coverage in the Technology hub. Start with What Is AI? for the model layer underneath voice-first interfaces, then Apple Intelligence and on-device AI for the compute budget on glasses, the humanoid renaissance for the shared perception stack, and the future of robotics. For the concept pages, see Platform Economics and Network Effects.
What are spatial computing interfaces?+
Spatial computing interfaces are user interfaces that exist in three-dimensional physical space rather than on a flat screen. Software windows, objects and controls are anchored to the room, and the user interacts through eye tracking, hand gestures, voice and physical accessories. Apple Vision Pro, Samsung Galaxy XR on Android XR, and Meta Quest headsets are the main consumer implementations in 2026.
What is the difference between spatial computing, mixed reality and augmented reality?+
Augmented reality overlays digital content on a live view of the world; mixed reality adds occlusion and physics so digital objects appear to sit behind or on real ones; spatial computing is the broader term for computing that understands and uses the geometry of physical space, including the input model. In practice the three terms describe the same device class from different angles: AR names the display effect, MR names the realism, spatial computing names the interaction model.
Which input method works best for spatial UI: gaze, gesture, voice or controllers?+
Each solves a different problem. Gaze plus pinch is fastest for selection at a distance and is the visionOS default. Voice is best for search and commands, and is the primary model Google built Android XR around with Gemini. Controllers remain best for precise, sustained input such as games, which is why visionOS 26 added PlayStation VR2 Sense controller support. Well-designed spatial apps offer at least two of these and never require one.
How do you optimize a spatial app for the visionOS App Store?+
Treat it as a search-only store. The app name (30 characters) and subtitle (30 characters) should carry the primary spatial keyword because there are no top charts to browse. Screenshots must be 3840 by 2160 pixels and should show the app in a real room, not a rendered void. Keep the video preview near 30 seconds and lead with the interaction, not the logo. Include the words users actually say — immersive, spatial, 3D — because voice search inside the headset uses natural language.
Are spatial computing apps a viable business in 2026?+
For headset-only apps the installed base is still small: IDC projects 3.2 million mixed reality headsets shipped in 2026 across all vendors. The larger opportunity is glasses — 13.6 million display-less and 3 million display glasses projected for 2026 — where the interface is voice and audio first. Apps that abstract their input layer so they run on both classes have the widest addressable market.
What is the Spatial Interface Stack?+
The Spatial Interface Stack is a five-layer model for spatial software: sensing (cameras, depth, eye and hand tracking), input grammar (gaze, pinch, voice, controllers), placement (where content sits and how it respects the room), persistence (whether content survives taking the device off), and distribution (how the app is found and installed). Each layer has a different owner — hardware vendor, OS, app developer, store — and design failures usually come from one layer assuming another's job.