For AI glasses OEMs, XR platform makers and ODMs
Eye tracking that fits the glasses you already designed. Under 10 milliwatts.
One camera and one LED per eye, in a 1.8 mm module. We look only through the pupil, never across the whole eye, so it mounts where your industrial design already has room.
Sub-degree accuracy, and no recalibration: not ever. Take the glasses off, put them back on, and it is still tracking.
Integration
Your frame does not have to change.
Immersix mounts a single module in the nasal corner of an ordinary spectacle frame. One camera, one LED, 1.8 × 1.8 × 2.2 mm per eye, over a standard MIPI serial interface.
No wide sightline across the eye. No ring of emitters around the lens. No thickened rim, no reshaped temple, nothing that forces an industrial designer to redraw the product to accommodate the sensor.
Applications
For AI glasses, gaze is the missing input.
The signal itself is simple: where the wearer is looking, right now. What it unlocks depends on what the device does with it: answer a question, save resources, or drive a display.
We supply the gaze stream; what you build on top of it is yours. Every one of these uses rests on the same two conditions: an estimate that stays right without asking the wearer to recalibrate, and a sensor small enough to be in the product at all.
AI glasses
Assistant context, resource efficiency, attention over timeResolving what "this" means
Stand at a supermarket shelf and ask "how much sugar is in this?" The scene camera sees forty products and has to guess. Gaze supplies the target; the voice supplies the verb.
Less power, lower latency
Running recognition over a full 4K frame is one of the largest power draws in a camera device. Gaze narrows it to the region being looked at — a fraction of the pixels, a fraction of the work.
A record of attention
Gaze is also a record over time: what held attention, what was skipped, what was returned to. Teams use it to time an interruption, or to test whether an interface works.
Additional uses with a display
AR glasses and headsets: rendering budget, focus cues, interactionFoveated rendering
Render at full resolution only where the eye is pointed. Published measurements report cutting more than 60% of the cost in the foveated portion of the pipeline, but only if the gaze estimate is right.
Correct focus cues
Where the eyes converge tells the display how far the content should sit, which is what varifocal optics need to place it at the right depth.
Gaze as the pointer
Look at a control, confirm with a pinch. Selecting small targets is an accuracy problem: at arm's length, a degree of error is about a centimeter of miss.
The problem
First-generation eye tracking needs an angle fashionable frames cannot provide, and looks at the wrong thing when it gets it.
Today's systems infer gaze from the outside of the eye: the edge of the pupil and glints reflected off the cornea. Catching those glints requires a camera with a wide, oblique view of the eye: set back far enough, or angled sharply enough, to see across the whole cornea. A headset has that room. A pair of glasses does not, because the optics sit at the lens plane, close to the eye and nearly edge-on, and a slim rim and temple leave nowhere to mount a camera that looks across the eye from a distance. The geometry PCCR requires is precisely the geometry a fashionable frame cannot give it.
There is a second problem, and it is why calibration exists at all. The visual axis (the line along which a person actually sees) is anchored inside the eye, in the retina. External features sit at an offset from it, and the offset is different for every person. A system watching the outside of the eye can only estimate the visual axis, and calibration is the process of producing that estimate. Every gaze figure such a system reports is an approximation carried forward from it.
Between them, those two facts produce every practical limitation below.
- Camera placementNo sightline across the cornea means no reliable glints. This, more than anything else, is why eye tracking has stayed in headsets rather than glasses.
- Power and real estateThe same approach needs 30+ frames per second and several emitters and cameras per eye. AI glasses have neither the budget nor the room.
- RecalibrationThe estimate drifts, and the user is asked to redo it, in some systems frequently.
- Slippage and device shiftsGlasses move on a face. When the device shifts relative to the eye, a pupil-based system loses its reference and has to be calibrated again.
- Population coverageAccuracy varies with eyelid shape and eye geometry, so performance differs across users in ways that are difficult to design around.
The technology
We image the retina through the pupil.
One module per eye, an infrared LED and a camera, captures images of the retina through the pupil. A one-time enrollment builds a map of that user's retinal features. From then on, every gaze estimate is made by matching what the camera sees against that map.
The visual axis is anchored in the retina. So tracking retinal features is not an improved estimate of where someone is looking: it tracks the visual axis itself. Where a pupil-and-glint system has to estimate that axis and then keep re-estimating it, we measure it. That is why there is nothing to recalibrate, and because the map is anchored to the retina rather than to the device, it does not care where the glasses are sitting on the face.
Enrolled once
One time, for the life of the user. No recalibration afterwards.
Anchored to the eye, not the frame
Device shift does not invalidate the reference, so slippage does not cost accuracy.
A view of the pupil is enough
We look through the pupil, focused at infinity, and never need the wide view of the whole eye.
Two engines, one camera
The retina engine is the accurate one, but it is too expensive to run on every frame. A lightweight pupil engine handles the frames in between: instant, and cheap enough to run continuously, with an accuracy that would slowly drift on its own. The retina engine re-anchors it before it can.
Retina engine
Expensive per frame, but rare enough that its cost per second stays small. Matches the live retinal image against the enrolled map and re-anchors the estimate.
Pupil engine
Instant and energy-efficient. Carries gaze between retina frames, with an accuracy that would gradually drift if left alone.
Both engines read from the same camera. Every gaze estimate comes from a single frame, with no averaging over time. This is what lets the system hold sub-degree accuracy at frame rates low enough to fit an all-day power budget.
Proof
See it tracking.
Gaze estimate overlaid on the scene camera, live. Two things to watch for:
- The accuracy on screen. The red diamond shows the gaze estimate.
- The moment the glasses come off and go back on. No calibration step follows. Tracking simply resumes.
Plays from YouTube. Watch on YouTube ↗
“The accuracy was definitely high, and it was always changing the color of the element I was looking at.”
“I also could move the glasses a bit and check that the system was able to work without any recalibration.”
Antony Vitillo (Skarredghost), The Ghost Howls. Independent hands-on, August 2026. Read the review ↗
Power
Retinal tracking works at a few frames per second.
That is what decides whether eye tracking ships in a pair of glasses or stays in a headset.
Retinal tracking re-anchors to a stable map, so it can run infrequently and still be right. This is what makes an all-day power budget possible.
System total, both eyes: compute, cameras and 940 nm illumination. Full breakdown in the specifications below.
Specifications
Target specifications.
| Metric | Value |
|---|---|
| Accuracy | < 0.5° @ p50 · < 1° @ p90 |
| Single lifetime enrollment time | 20 seconds |
| Field of view | 71° diagonal (55° × 45°) |
| Camera module | 1.8 × 1.8 × 2.2 mm |
| Max update rate | 120 fps |
| Ambient light | 50,000 lux (~10,000 lux at the eye after AR lens filtering) |
Scroll sideways for the remaining columns →
| Component | @ 1 fps | @ 10 fps | @ 100 fps |
|---|---|---|---|
| Compute (Alif M55 + hardware accelerator) | ~1.7 mW | ~2.6 mW | ~14.9 mW |
| Cameras (one per eye) | ~5.6 mW | ~6.3 mW | ~13.2 mW |
| Illumination (940 nm IR LED, one per eye) | ~0.1 mW | ~0.6 mW | ~6.4 mW |
| Total | ~7.3 mW | ~9.5 mW | ~34.5 mW |
Scroll sideways for the remaining columns →
Power figures are system totals for both eyes.
Download the full data sheet (PDF)Built-in capability: Authentication
The sensor that tracks your gaze already knows who you are.
A retina is a biometric, richer than a fingerprint and highly resistant to presentation attacks. The optical constraints that make retinal imaging difficult also make it hard to spoof: the image forms through the pupil, off the interior of the eye, focused at infinity. There is no photograph and no generated image you can hold up to a camera that reproduces one.
Enrollment already builds a retinal map, so authentication costs nothing extra in hardware or in user effort.
No separate step
The device recognizes the wearer as soon as it is on, and a shared headset loads the right profile by itself. The same recognition is what makes payments, enterprise login and age verification possible without a separate authentication step.
Resistant to presentation attacks
The features are inside the eye and never visible externally, so there is no photograph, video or generated image to copy them from, and the image only forms through a live pupil, focused at infinity.
Unique and stable for life
A retinal feature map is unique to the individual and stable for the life of the user. The enrollment that makes tracking work is the same one that makes identity work, so neither needs redoing.
Continuous and passive
Ongoing confirmation that the person wearing the device is still the person who unlocked it, not a check at the door and then nothing.
Test it on your own users.
Evaluation kits are available. A kit is a wearable rig and a laptop running our capture and analysis software: you enroll your own subjects, run your own protocol, and export frame-by-frame accuracy, precision and trackability data to analyze yourself. No filtering, outliers retained. We would rather you measured us than believed us.
The rig is engineering hardware and looks like it: the camera modules on it are 3.3 mm rather than the 1.8 mm production part, and compute runs on the laptop rather than on the frame. We send it because the technology is the thing worth evaluating.
Team
Built by people who have shipped optics, silicon and computer vision.

Algorithms and computer vision expert. Developed a novel positional tracking technology and a new method of position coding. Since late 2016 Ori has focused his attention on solving the gaze tracking problem. Ph.D. candidate. M.Sc. summa cum laude, B.Sc. magna cum laude.
LinkedIn ↗
More than 25 years in the high-tech industry, leading startup companies and R&D organizations developing innovative products for new markets. Previously CEO of FarmSee, CEO of PointGrab, and CTO of N-trig. B.Sc. in Electrical Engineering, cum laude, Tel Aviv University.
LinkedIn ↗
Over 28 years of technological experience growing startups from inception to exit. Former director of hardware at Modu. Builds and manages multidisciplinary teams across hardware, software, mechanical and industrial design, with a focus on consumer device development.
LinkedIn ↗
Over 20 years as a serial entrepreneur and investor focusing on AI, fintech and computer vision. Founded Fraud Sciences (acquired by PayPal), ClarityRay (acquired by Yahoo) and Trivnet (acquired by Gemalto). First investor in Moon Active, Supersonic (acquired by Unity) and Crosswise (acquired by Oracle).
LinkedIn ↗