Repository navigation
Releases: amariichi/StereoSplatViewer
Release list
v1.4.0
Two changes to the desktop editor, both about photographs with something far away in them.
The lens can be changed after you have seen the scene
The lens a photograph was taken with decides the shape of the reconstruction, not merely its scale, and a photograph that does not record its own is unprojected through a 30 mm default that is usually wrong. The editor asked for the number before uploading — which is advice nobody can follow, because the whole difficulty is that it cannot be judged until the scene exists and you have looked at it.
Rebuild at this lens, under the same field, builds the photograph the server is already holding again through a new number. Nothing is uploaded a second time: the picture is still in the job directory, and the scene keeps its place, so this can be tried as often as it takes to find the right one.
The backend could always do this and the phone could always ask for it — but only for a photograph that recorded no lens of its own. The desktop, which is the one place with a screen large enough to judge the shape on, could set the number once and never revise it. It is now offered for any held scene, including one whose EXIF recorded a lens, because a photograph can record one and record it wrongly. A 360 is six reconstructions merged and cannot be rebuilt this way.
It is a separate button from Generate deliberately. Generate needs a file, clears the stored scene and starts a different one; this takes no file and replaces the scene where it stands.
Turning a scene with a distant background is controllable
Dragging to turn a photograph with a background — a landscape, a room with a window, anything with sky — threw the subject across the frame faster than a mouse could be aimed.
The camera orbits a point at the median distance of the gaussians. That is right when the subject fills the frame, and wrong when it does not, because the median lands in whichever half is larger: in a landscape, the background. In one measured scene the nearest splat sat at 1.20 units and the median at 57.83, so the camera swung on a radius forty-eight times the distance to the subject, and a ten pixel drag moved it three times further than the subject was away.
Finding the subject instead was tried and abandoned against that same scene. It is not two clusters with a gap between them but four humps with no gap at all: the largest step between neighbouring sampled distances across its middle nine tenths is two and a half per cent, which is noise, and a valley in log space exists but there are two of them, five times apart.
So the pivot is bounded rather than guessed. What makes a distant pivot unusable is not how far away it is but how that compares to the nearest thing on screen, because that ratio is exactly how much a drag is magnified there. A scene whose subject fills the frame is left alone; so is a landscape with nothing near it, where everything is far and the ratio is small.
In stereo this counted twice, since pivot mode puts the plane of the screen at the same distance and had been putting it on the horizon.
Nothing on the phone changes.
The full description of what this project is, and the two external conditions worth reading before installing, are in v1.0.
v1.3.2
One fix, for anyone who turns the phone while looking through it.
Turning the phone now works with Hold level off. Rotating the phone about its upright axis — bringing the left edge toward you and pushing the right edge away — showed the wrong side of the scene: the far edge of the splat and the darkness past it, where the near edge belonged. Turning Hold level on fixed it, and turning it off brought it back.
Which way the glass is turned is not a stabiliser. It is what makes the window a window. A camera watching your face cannot tell a turned phone from a moved head, and the two ask for opposite corrections, so without the phone's own sense of its heading the scene swings across its whole depth — far enough, at twenty degrees, to slide the background nearly two half-frames past the subject. That measurement existed and was correct. It was simply unreachable, because switching Hold level off stopped the sensors outright.
The sensors now run for as long as the camera does. Hold level keeps the name and the button, and now means only what it says: holding the model's horizon level against the phone's pitch and roll, up to eighteen degrees. Turning the phone is corrected either way. Switching the button on also adopts the posture the phone is in at that moment, rather than whichever one the sensors happened to start in, so it no longer matters how long ago you pressed Start 3D.
If motion or orientation access was refused, nothing here changes: the view still works, and a turned phone is still read as a moved head.
Nothing else changes.
The full description of what this project is, and the two external conditions worth reading before installing, are in v1.0. What changed in True Window and in head tracking is in v1.3.0.
v1.3.1
One fix, for anyone who uses the phone viewer in both orientations.
Hold level now survives turning the phone. Rotating between portrait and landscape left it switched on but reading the new screen frame through everything it had learned in the old one, and no amount of holding the phone still would settle the stale correction. Turning Hold level off and on cleared it, because that discarded the whole sensor tracker and built a fresh one. Nothing short of that did.
Only what the sensors had measured was ever wrong. The tracker now clears its filters and reference postures in place when the orientation changes, keeping the motion and orientation permissions that iOS grants only from a button press, and keeping both event subscriptions alive. Hold level stays visibly on, nothing asks for permission a second time, and the previous orientation's correction disappears the moment the layout turns.
The screen changes orientation before the hand has finished the quarter turn, so the first sample after it is usually taken mid-turn. Adopting that one would define level as an angle nobody meant, or push the correction straight to its 18-degree limit. A quarter of a second of samples is discarded instead, and the first settled reading becomes the new reference — about as long as the turn itself, and not a wait you can feel.
Head tracking, the scene on screen, the eye-distance calibration, pinch framing, and True Window's physical scale all carry across the turn untouched. Both directions, in True Window and in photo mode.
Also in this release, the README described two ways to serve the viewer over HTTPS. There are three. If tailscale serve is already mapped to the dev port on your machine, the best answer is scripts/dev.sh with no option at all: Tailscale terminates TLS with its own genuine certificate, and the phone opens the address with no port number and no warning. Adding --https or --tailscale there makes the dev server speak HTTPS while serve still forwards plain HTTP to it, and the mismatch answers 502. The script had always detected this and said so; the README had not.
Nothing else changes. If v1.3.0 works for you and you hold the phone one way, there is nothing here you need.
The full description of what this project is, and the two external conditions worth reading before installing, are in v1.0. What changed in True Window and in head tracking is in v1.3.0.
v1.3.0
Highlights
- True Window now treats the phone display as a fixed physical aperture and scales the miniature independently.
- Metric cyclopean-eye tracking, live panel geometry, and optional screen-size calibration keep the observer and rendered world in the same scale.
- Hold level now smooths pitch, roll, and relative phone yaw while preserving real head translation.
- Embedded PLY camera intrinsics are preferred when available, with the existing splat-distribution fallback retained.
Compatibility
- Photo-preserving mode keeps its source-frame behavior.
- If orientation access is unavailable, gravity pitch/roll levelling continues and only phone-yaw separation is omitted.
v1.2.0
The phone can now be told what lens took a photograph. If your pictures carry that in their EXIF, nothing here will ever ask you anything and you can skip to the last line.
Why it matters more than it sounds. ml-sharp reads the focal length from the file and assumes 30 mm when it finds none. That assumption is not about size: it is the field of view the scene is unprojected through, so it sets the shape. The same photograph built at three lenses puts its subject 2.5, 4.1 and 6.9 metres away, with the depth stretched to match. A portrait taken at 85 mm and built at 30 comes out pressed flat — recognisably the right picture, with the depth squeezed out of it.
The editor has had a field for this. The phone, which is the only place you can paste a picture, had none.
-
It asks once, before building. A photograph that records its own lens is built straight away. One that does not is held, and the field appears; building starts when you answer. Leave it blank to accept the 30 mm that would have been assumed anyway. A portrait is usually 50 to 85; a phone's own camera is around 24 to 28.
-
Full-width digits are read as ordinary ones, so there is no need to change input mode to type two characters.
-
You can change your mind. The lens cannot be judged before the scene exists, so the field stays: type a different number and press Rebuild. That runs
ml-sharpagain and takes about as long as the first time. It runs only when you press it. A 360 scene cannot be rebuilt this way, being six reconstructions merged rather than one. -
Tipping the device now moves the view the other way round. Reported as backwards, and it was.
One note for anyone who had already made a scene: scenes built before this release carry no record of what lens they used, and are left alone rather than guessed about. Paste the picture again if you want to set one.
The full description of what this project is, and the two external conditions worth reading before installing, are in v1.0. What changed in the phone viewer itself is in v1.1.0.
v1.1.1
One fix, for anyone using the phone viewer away from a fast connection.
A scene is no longer downloaded twice. A phone that goes into standby has its tab evicted, and coming back meant fetching the whole scene again — eleven megabytes, over whatever connection happens to be there. The file was cacheable all along and nothing said so, which left the decision to whatever each browser inferred. It is now stated outright: a scene never changes once it exists, so it may be kept for a year. Returning from standby should now show the picture straight away rather than the download bar.
Also fixed, and the half more likely to have caused trouble later: the address the viewer polls to notice a new scene is now explicitly not cacheable. It never was cached in practice, but nothing had said it must not be, and a browser that decided otherwise would have pinned the page to whichever scene was current when it first asked — with no scene made afterwards ever appearing.
Nothing else changes. If v1.1.0 works for you on a fast local network, there is nothing here you need.
The full description of what this project is, and the two external conditions worth reading before installing, are in v1.0. What changed in the phone viewer itself is in v1.1.0.
v1.1.0
The phone viewer changes in this one, in a way you will notice. If you have used it before, read the first two lines.
The controls have changed. One finger now turns the miniature — sideways spins it, up and down tips it, and a diagonal does both at once. Two fingers pan and pinch. Turning two fingers against each other no longer does anything; a hand cannot twist very far, and that gesture was sharing the pair with the pinch.
The picture is now drawn from where your eye actually is. That is what the True window button says, and it is on by default; the old behaviour is still there if you turn it off.
The difference: a face used to stretch, sometimes badly, when the phone was tilted steeply or the picture was dragged toward an edge. The projection puts a distortion into the picture on purpose, and looking at the screen from the angle it was built for is what cancels it. It was built for the photograph's angle, which is wide — a phone at arm's length spans a small fraction of it — so most of the distortion survived, worst where the angle is largest, which is the edges.
What you give up is that the whole frame is no longer on screen at every depth. Content on the glass is all there; the further back it sits, the more the window crops it. That is what a window does, and it is why the far background is now a narrower slice than the photograph contains.
-
Scenes arrive much faster on mobile data. The phone is sent a compressed copy, about a sixth the size — 66 MB became 11 MB on the scenes tested — and the page now shows how far through the download it is instead of sitting black. The editor still uses the full file, so exported images are unaffected. Made with
splat-transform, which already comes with the frontend, so there is nothing extra to install; if it cannot run, the phone is sent the full file and the job log says why. -
The phone follows the editor. Make a scene on the desktop and it appears within a few seconds, with no reload. An address that names a particular scene still wins.
-
Photographs wider than the screen are no longer cut off. The scene is fitted to the frame's height, so an upright phone was losing the sides — about a third of the width for a landscape photograph. The page now opens zoomed out far enough to show the whole width, with bars above and below. Pinching in from there works as before.
-
The distance calibration is more accurate. I am at N mm corrects a scale error, and the correction was only being applied to the distance. It now applies to sideways and vertical position too, which is what decides how far off the axis your eye is.
-
A notice can be dismissed by tapping it. Several of them were not failures at all — refused motion access, for one — and they stayed on screen for the rest of the session.
-
Depth is a new button. It slides the miniature further behind the glass, which settles how far back it sits and how widely it swings when you turn it with a finger. It is not a depth-strength control: moving a whole scene away flattens it rather than deepening it.
Four limitations are known and left as they are: a compressed scene that arrives corrupt is not retried as the full file; replacing a scene mid-download cannot stop the download, because the renderer offers no way to abandon one; ?job= without name is not treated as naming a scene; and the fallback head-tracking path keeps its own limit on reported position.
The full description of what this project is, and the two external conditions worth reading before installing, are in v1.0.
v1.0.2
Housekeeping, with one thing worth knowing if this shares a GPU with anything else. Nothing here changes what the app does — the scenes you make and the way you look at them are identical to v1.0.1, and the PLY files come out byte for byte the same.
-
The renderer is now PlayCanvas 2.21.4, up from 2.14.4. It was pinned because the newer engine rendered a blank canvas: unified rendering became the default at 2.15, and the two things this viewer reached for through an undocumented sorter — the cue to draw, and the gaussian positions it frames a scene from — both stopped being reachable that way. Both are published API now, and the cue is read from the engine's own constant, so a rename upstream becomes a build error rather than a blank canvas. The pin existed to buy time, and non-unified rendering is documented as going away, so this was when rather than whether.
A second break came with it, and was not in anyone's release notes: 2.21.4 chooses how to read a scene from the file extension of the name it is given, and "open a .ply from this machine" hands the viewer a
blob:address, which has no name at all. Scenes from the server were never affected, so checking only that route would have shipped a viewer that could not open a local file.Stereo is unchanged and was measured rather than assumed: the two eyes still carry projection matrices identical in every vertical term and opposite in the horizontal skew — not within a pixel, the same bits. Drawing still happens only when something moved: with a scene loaded and nothing touched, 180 frames pass and none are drawn.
-
The scene generator gives the GPU back between images. The resident ml-sharp worker keeps the model in memory so that each scene costs about 3 seconds instead of 16.5, and that model is 2.7 GB. It was also keeping one image's working memory for the rest of its life, so an idle worker held 11.5 GB. It now holds 3.1 GB.
This costs nothing: the same image three times took 2.63 s with the memory kept, then 2.38 s and 2.37 s with it released after each. If you have been running this on a card where 11.5 GB was awkward, that constraint is gone.
Verified by generating a scene and looking at it on a phone, which is the only check that catches this kind of fault — the upgrade passed typecheck, lint, all 142 tests and the build at a point when it was rendering nothing at all.
The full description of what this project is, and the two external conditions worth reading before installing, are in v1.0.
v1.0.1
Housekeeping. Nothing here changes what the app does — if v1.0 works for you, there is no reason to hurry.
- CI runs once per commit rather than twice. A branch and the pull request for it are the same commit, and the workflow triggered on both, so every pull request showed two rows for each check. It now runs on pull requests, on
main, and on tags. Superseded runs are cancelled too. - The backend dependencies are locked in
backend/uv.lock, and installed from it.pyproject.tomlgave version ranges and nothing recorded what they resolved to, so two machines installing on different days could get different versions of numpy or pillow with no way to tell afterwards.scripts/setup_wsl.shinstalls from the lock whenuvis present and says plainly when it is falling back to the ranges. CI checks the lock andpyproject.tomlagree, and fails when they do not.
After changing a backend dependency, run uv lock --project backend and commit the result.
The full description of what this project is, and the two external conditions worth reading before installing, are in v1.0.
v1.0
The first release that does not carry a fork of the SuperSplat editor.
What you can do that you could not before
Look around a scene by moving, on a phone. /viewer.html uses the front camera to find where your head is and draws the scene from that point, so the screen behaves like a window onto something standing just behind the glass. Moving your head is what reveals the shape. One finger slides the scene within the frame, two fingers zoom, twist to turn it, drag together to tip it up to 55°.
Free-view a pair cross-eyed. Side-by-side output was parallel-view only, which cannot be fused on a display wider than your eyes. There is now a swap, and exports follow it.
Get a scene in about three and a half seconds instead of sixteen and a half. ml-sharp is kept warm in one process rather than started per image, for byte-identical output. A 360 scene went from about a hundred seconds to thirty.
What was wrong and is now fixed
- The stereo was toe-in — two cameras rotated inward, which introduces vertical disparity growing toward the corners that eyes cannot fuse comfortably. It is off-axis now, measured at exactly zero vertical disparity.
- 360 scenes were different every time. They were recentred on a bounding box measured from a random subsample, which also moved the capture point away from the eye. Two runs of one image gave centres of z=-20.1 and z=-4.6.
- A 360 job with no merge tool reported success and showed nothing, with the explanation rendered inside the section that never appeared.
- Generated scenes were deleted whenever the backend restarted. Since each upload already clears the data root, that only ever destroyed work.
- Opening a local .ply deleted whatever was on the server — including whatever a phone was looking at — although the file never reaches the server at all.
- Three stereo controls did nothing: Framing lock, Comfort lock, Comfort strength. The zero-parallax "Double click" mode never responded to a double click either; it is now Fixed distance, with a field.
New
- An optional treatment for a sky overexposed to flat white, where the depth model has nothing to measure and a quarter of the sky can be placed at the subject's own distance — a white wall around the head. Measured on one photograph: 25.7% → 2.9%. Off by default, because it helps where that is the fault and mildly hurts where it is not.
- The editor is laid out around the preview rather than as numbered panels.
- CI on every branch and pull request.
Before you install
Two conditions are set by other people, not by this project, and both are in docs/THIRD_PARTY_NOTICES.md:
- The ml-sharp model weights are released for research — "for the sole purpose of scientific research of artificial intelligence and machine-learning technology". Read
LICENSE_MODELbefore using the output for anything else. - The phone viewer downloads MediaPipe from a CDN when it starts, so that one page needs an internet connection and tells two hosts it was opened. Everything else runs on your machine.
Node 20.19+ is required (22 or 24 recommended). ml-sharp is installed separately; nothing is vendored here.