3D Photo Converter
Turn one ordinary photo into something you can see in 3D. A depth model reads the picture and a second viewpoint is built from it, which then becomes a side-by-side pair for free-viewing, a red-cyan anaglyph for coloured glasses, a wiggle GIF that needs neither, or the depth map itself. Everything runs in your tab; the photo is never uploaded.
- Four outputs from one depth pass — side by side, anaglyph, wiggle animation, and depth map — with every control updating the preview as you drag it
- The wiggle needs no glasses and no free-viewing knack, so it is the one output that asks nothing at all of whoever you show it to
Drop your photo here
It stays on your machine and is never uploaded
Reads JPEG, PNG, WebP, AVIF, GIF, and BMP; an animated GIF contributes its first frame. Rotation recorded by the camera is applied, so a portrait shot arrives upright.
Add a photo to beginAny ordinary 2D photo works. The depth is worked out from the picture itself, so no special camera is needed.
Overview
What the tool does with a flat photo, and why each part is there.
- 01
Depth from a single photograph
A neural depth model looks at the image and estimates relative near-to-far depth for every pixel — not a measurement of distance, but an ordering of what sits in front of what. It reads the same cues a person reads a photograph with: occlusion, perspective, texture gradient, relative size. That is what makes an ordinary snapshot usable as a source rather than only a shot taken on a stereo camera.
- 02
A synthesised second viewpoint
The depth map drives a horizontal shift: near pixels move further than distant ones, which is exactly what separates the two images your eyes receive in the real world. Both views are shifted by half the total rather than holding one still, so the framing stays centred and each view opens only half as much hidden area.
- 03
Four outputs from one depth pass
The model runs once per photo and everything after it is arithmetic, so side by side, anaglyph, wiggle, and depth map all come out of the same estimate at no extra cost. That is also why every slider updates the picture as you drag it, instead of asking you to re-run and wait.
- 04
Anaglyph done properly
Seven methods, including Dubois' least-squares projection for both red-cyan and green-magenta. Rather than assigning channels to eyes, it solves for the pair of images that look closest to the original once each filter's leakage is accounted for, which is why reds stop turning into black holes. A colour-strength control and an eye swap cover the two things that most often go wrong.
- 05
A wiggle that needs nothing at all
The animation rocks between viewpoints, so the movement supplies the parallax that two eyes normally would. No glasses, no learned knack, and it survives being posted somewhere. The loop is built so its turning points are not duplicated, which is what stops the once-a-loop stutter a naive sweep produces, and the edges the outermost views had to invent are trimmed off.
- 06
Depth edges snapped to the picture
The model works at around 518 pixels, so its depth boundaries arrive soft and slightly off the object, which is what puts a halo of background around a subject. An edge-aware filter re-averages depth using the photograph's own edges as the guide, so depth stops crossing a boundary the picture already knows about.
- 07
Gap filling that uses the background
Shifting pixels sideways exposes areas the camera never recorded, always at the edges of nearer objects. Each gap is patched from whichever side is further away, because that is what the missing area belongs to, and the background is reflected into the gap so texture carries across rather than smearing into a flat streak.
- 08
The size question answered on screen
Whether a parallel pair can be fused depends on how wide it is drawn, not on its pixel dimensions, and that catches nearly everyone. The preview has a size control with an estimate of the physical width of each half, plus a button that puts it inside the roughly 63 mm limit straight away.
How to use
From a flat photo to something you can see in 3D.
- 01
Add a photo. Any ordinary 2D image works — a phone snapshot, a scan, a download. Nothing leaves your browser.
- 02
Press Generate 3D. The depth model is fetched the first time and then reads the photo, which takes a few seconds.
- 03
Pick what you want to make. Choose Wiggle for something easy to share; anaglyph if you have coloured glasses; side by side for free-viewing, or side by side in Parallel order for a headset; depth map if another program is going to use it.
- 04
Adjust the 3D strength until depth reads clearly without smearing at the edges of objects. The preview updates as you drag.
- 05
Set the screen plane. For side by side leave it at zero so everything floats towards you; for anaglyph and wiggle put it on your subject.
- 06
Check the depth map view if a result disappoints. Bright is near, dark is far. A misread scene shows up here far more plainly than in the result.
- 07
For side by side, use the size control under the preview to shrink the pair until each half is around 63 mm, then relax or cross your eyes.
- 08
Press Download. Still outputs save as PNG, JPEG, or WebP; the wiggle is encoded as an animated GIF at that point.
Details
The details that decide whether the result is something you can actually see.
- Monocular depth estimation on the photo itself, so an ordinary snapshot is a valid source with no stereo camera involved
- Side-by-side output in parallel or cross-eyed order; VR photo viewers generally expect the Parallel order
- Seven anaglyph methods, including Dubois least-squares projection for red-cyan and for green-magenta
- A colour-strength control and a left/right swap, covering the two commonest reasons an anaglyph refuses to fuse
- Animated wiggle output as a GIF that reads as 3D without glasses and without any free-viewing skill
- Wiggle frame count, speed, and motion curve, with a loop whose turning points are not duplicated so it does not stutter
- Automatic trimming of the invented edge strip in wiggle mode, where a changing patch would otherwise draw the eye
- Depth map export in greyscale or three colour ramps, with an invert switch for software that expects near to be dark
- Edge-aware depth refinement guided by the photograph, which removes most of the halo around subjects
- Both views shifted by half the total, so the framing stays centred and each view opens only half the hidden area
- Gaps patched from the far side of each edge and mirrored inward, rather than smeared out from the foreground
- Live preview: the depth model runs once per photo, and every control after that redraws immediately
- A display-size control with an estimated physical width, which is the actual answer to "why can I not see it"
- PNG, JPEG, or WebP for stills with a quality control, and an animated GIF for the wiggle
- Everything runs in the browser tab: no upload, no queue, no account, and no watermark
Use cases
Why people want a flat photo turned into a stereo image.
-
Making a wigglegram to post
The wiggle is the 3D format that is easiest to share in a feed: no glasses and nothing to learn. Multi-lens film cameras used to be the only way to make one, and they needed the shot planned in advance. This makes one from a photo you already have.
-
Old photographs and scans
Family pictures, archive prints, and anything shot before stereo capture existed are flat by nature and cannot be re-taken. Depth estimation is the only route to a second viewpoint for them, and the result can be looked at without buying anything.
-
Anaglyphs for a room full of people
Free-viewing is a solo skill that not everyone has. A packet of cardboard red-cyan glasses costs almost nothing and lets a whole room see the same picture at once, which is why anaglyph is still the format for a classroom, a talk, or a printed page.
-
Depth maps for other software
Parallax web effects, displacement in a 3D package, focus and relighting in a compositor, and depth-conditioned image generation all want a depth map and rarely provide one. Exporting it here is often the whole reason to open the tool.
-
Photos for a VR headset
Headset photo viewers commonly support side-by-side images. Choose Parallel in this tool, then set side by side and the left/right eye order in the viewer if it asks: the file carries no stereo metadata. The same pair can also be free-viewed on a phone.
-
Landscapes, portraits, and product shots
Scenes with a clear foreground and a real background convert far better than flat or cluttered ones. A landscape with layered distance, a person against a room, or a single object on a plain surface are close to the ideal case: unambiguous depth structure, clean edges, and little to invent behind them.
-
Teaching how stereo vision works
Being able to convert a photo and then put the depth map beside it makes the mechanism visible in a way a diagram does not. The failure cases are as instructive as the successes, and switching between anaglyph, side by side, and wiggle shows three different answers to the same problem.
-
Pictures that have to stay on your machine
Family photographs, client work, and anything under an agreement have no business on someone else's server for the sake of an experiment. Processing stays in the tab: only the depth model is downloaded, while the photo itself never leaves your device.
See also
The depth model reads whatever framing you give it, so crop and straighten before converting rather than after, which is a job for Image Cropper. Very large photographs cost decode memory without improving the depth estimate, and cutting one down to size first is what Image Resizer is for. The same conversion applied to footage rather than a single frame, with segment-by-segment processing and the audio carried across, is Stereo 3D Converter.
How glasses-free 3D actually works
The mechanism behind each output, and the constraints that decide whether you can see it.
-
Two pictures, one for each eye
All stereo 3D works the same way: show each eye a slightly different view of the same scene and the brain fuses them into depth. Cinemas use polarised glasses to route the two images; a headset uses two screens; an anaglyph uses colour. Free-viewing skips the hardware entirely by having you aim your eyes so each lands on its own half of a side-by-side pair.
-
Parallel versus cross-eyed
These are two ways of aiming. In parallel viewing your eyes relax outward as though looking past the screen, and the left image belongs on the left. Cross-eyed viewing points them inward, in front of the screen, so the halves are swapped. The same file cannot serve both — viewed the wrong way, depth inverts and the picture turns inside out while still looking convincingly three-dimensional.
-
Why size matters for parallel viewing
To fuse a parallel pair your eyes must rotate no further apart than straight ahead, because human eyes cannot diverge. That caps how far apart matching points can sit at roughly 63 mm, the distance between your pupils. On a phone that is easy. On a large monitor showing a full-width pair, each half is tens of centimetres across and fusing it is simply not possible, whatever the resolution.
-
What an anaglyph trades away
Colour filters can separate two images because each lets through what the other blocks — but no real filter is perfect, so some of each view leaks into the wrong eye as a ghost. Worse, a saturated red object is bright to one eye and black to the other, and the brain refuses to fuse the two: retinal rivalry, and the main reason cheap anaglyphs are uncomfortable. Every method here is a different answer to that trade, from keeping all the colour to throwing all of it away.
-
Why the wiggle works without anything
Stereo is not the only depth cue you have. Motion parallax — nearer things sweeping across your view faster than distant ones — works with one eye closed, and is why leaning sideways tells you what is in front of what. An animation that rocks between two viewpoints hands you that cue directly, which is why a wiggle reads as depth to people who cannot free-view and are not wearing anything.
-
Where the artifacts come from
Two things limit any 2D to 3D conversion built on estimated depth. The model can misread a scene — a poster read as a window, a reflection read as distance — and no setting recovers from that, which is why the depth map is worth a look. And moving pixels sideways exposes background that was never photographed, so it has to be invented. Both get worse as strength rises, which is why modest settings usually look better.
Best practices
Habits that get a result worth looking at.
- Start with the wiggle if you are showing somebody else, since it is the only output that needs nothing from them
- Keep the 3D strength modest and raise it only until depth reads clearly — the artifacts grow faster than the effect does
- Look at the depth map when a result disappoints, because a misread scene cannot be fixed by any slider and tells you immediately to try a different photo
- Put the screen plane on your subject for anaglyph and wiggle, and leave it at zero for parallel free-viewing
- Try Dubois before any other anaglyph method, and reach for the colour-strength control rather than a different method when strong reds shimmer
- Swap the eyes if an anaglyph looks three-dimensional but wrong, since glasses get worn the other way round more often than anyone admits
- Use the display-size control for parallel pairs instead of assuming the effect is broken, because width on screen is what decides whether it fuses
- Try cross-eyed first if free-viewing is new to you, then switch to parallel once you can fuse an image reliably
- Prefer photographs with a clear foreground and a real background — flat scenes, tight close-ups, and heavy motion blur all give the model little to work with
- Keep wiggle output small: a GIF stores every frame in full, so width costs far more here than it does on a still
- Save depth maps as PNG, since compression artifacts in a depth map become wrong geometry as soon as another program reads it
- Crop and straighten before converting rather than after, so the depth model sees the framing you actually intend to use
Limitations
What this tool cannot do, and where the results will disappoint.
- The second viewpoint is inferred, not photographed. Depth is estimated from a single image, and where that estimate is wrong the geometry is wrong with it — this is not equivalent to a shot taken on a stereo camera.
- The model misreads some scenes. Reflections, glass, printed images of scenes, flat surfaces with strong texture, and heavy motion blur are all read unreliably. The depth map view exists so you can see it rather than guess.
- Areas hidden behind objects have to be invented. Nothing photographed them, so the fill is a reflection of neighbouring background. It is unobtrusive at modest strength and obvious at high strength, particularly around thin objects and hair.
- Free-viewing is a learned skill. A significant number of people cannot do it at first, and some never find it comfortable. Cross-eyed is easier for most; the anaglyph and the wiggle remove the problem entirely.
- Parallel viewing is limited by the width of the image on your screen, not by its pixel dimensions. A pair that fuses on a phone can be impossible on a monitor at full size, and no setting here changes that — only how large you display it.
- Anaglyphs always cost colour accuracy, and no method removes ghosting completely. Strong reds and greens are the hardest cases, and a picture that is mostly one saturated hue may never look right in any anaglyph method.
- The wiggle is a GIF, so it is limited to 256 colours and gets large quickly. Dithering hides the banding but adds noise that differs frame to frame, and neither choice is free.
- Depth is relative, not measured. Values say what is in front of what within this one picture, with no absolute scale, and two photos of the same scene will not agree on their numbers.
- The depth map is 8-bit, so it holds 256 distinguishable levels. That is plenty for parallax and masking and not enough for precise displacement work.
- It is not the anamorphic illusion on curved outdoor LED billboards. That is built for one specific screen geometry and viewing position and cannot be derived from a flat photo, whatever a given app claims.
- Metadata is not carried over. The output is drawn on a canvas, so EXIF, colour profiles, and GPS tags are all left behind — which is convenient for privacy and a loss if you wanted them.
- Wide-gamut and HDR photos are converted through the browser's ordinary canvas, so they come out in standard dynamic range and sRGB.
- The depth model is 26 to 47 MB on first use, depending on whether your browser has WebGPU. It is cached afterwards, but a fresh browser profile carries the download again.
- One photo at a time. There is no batch mode, because the settings that make a good result are judgement calls about a particular picture rather than something to apply blindly to a folder.
FAQ
Questions that come up when turning photos into 3D images.
How do I turn a 2D photo into a 3D image?
Add the photo and press Generate 3D. A depth model reads the picture to work out what is near and what is far, and a second viewpoint is built by shifting pixels sideways in proportion to their depth. Then choose what to make from it: a side-by-side pair you view by aiming your eyes, a red-cyan anaglyph for coloured glasses, an animated wiggle that needs neither, or the depth map itself. The model runs once, so after that every control updates the preview as you drag it and you can try all four without waiting.
Do I need 3D glasses?
Only for the anaglyph. The wiggle output needs nothing at all — it animates between the two viewpoints, and that movement gives your brain the same depth cue that leaning sideways does, which is why it still reads as depth for someone who sees with one eye only. The side-by-side pair also needs no hardware, but it does need free-viewing, which is a knack some people pick up in minutes and others never find comfortable. If you want to hand a 3D photo to somebody else with no instructions, use the wiggle.
How do I make a wigglegram from a photo?
Choose the Wiggle output after generating depth. A wigglegram is an animation that rocks between slightly different viewpoints of the same scene, and it originally required a film camera with several lenses set side by side. Here the extra viewpoints are synthesised from the depth map instead, so one ordinary photo is enough. Set the frame count — two gives the classic hard alternation, more gives smoother motion — pick a speed around 10 to 14 frames a second, and download an animated GIF. The preview animates live at whatever settings you choose, so you can judge it before encoding anything.
What is the best anaglyph method for red-cyan glasses?
Dubois, in almost every case. The obvious approach is to give the red channel to one eye and green and blue to the other, but real filters leak, and a saturated red object then appears bright to one eye and black to the other — the brain refuses to fuse that, which is what makes cheap anaglyphs uncomfortable. Dubois instead solves for the pair of images whose appearance through actual filters lands closest to the original, so reds survive and ghosting largely disappears. If a picture still shimmers, lower the colour strength before trying another method; greyscale and monochrome remove the problem entirely at the cost of all colour.
Is my photo uploaded anywhere?
No. The depth model runs through WebGPU or WebAssembly inside your browser tab, and every other step is arithmetic on a canvas. Your photo is read from disk into the tab's memory and the finished image is written back out as a download. The depth model itself is downloaded from the network the first time you use the tool, but the photo and generated result are never uploaded.
Why does the 3D look wrong or inside out?
Three common causes, in order of likelihood. First, the glasses are on the other way round, or the side-by-side pair is being viewed with the wrong technique — both invert the depth, and an inverted picture still looks convincingly three-dimensional, so it is easy to miss. Use the swap control for anaglyphs, or switch between parallel and cross-eyed. Second, the depth model misread the scene, which the depth map view will show you immediately: reflections, glass, and printed pictures within the shot are the usual culprits, and no setting recovers from a wrong depth map. Third, the strength is simply too high, which reads as smearing around objects rather than as depth.
Can I get a depth map from a photo?
Yes — choose the Depth map output and download it. For another program, select Greyscale and PNG: that gives you an 8-bit map at the output size you picked, aligned pixel for pixel with the scaled photo, with bright meaning near by default and an invert switch for software that expects the opposite. JPEG, WebP, and the three colour ramps are for sharing or inspecting the map, not geometry. It is relative depth: it tells you what is in front of what within that one picture, not how far away anything is in metres. That is what parallax effects, displacement, distance-based masking, and depth-conditioned image generation need.
What is the difference between parallel and cross-eyed viewing?
They are two ways of aiming your eyes, and each needs the halves in a different order. Parallel viewing relaxes the eyes as though looking through the screen, and wants the left eye's image on the left. Cross-eyed viewing points them inward, in front of the screen, so the halves are swapped. Most people find cross-eyed easier to learn. It avoids parallel viewing's hard divergence ceiling, but a moderate display size is still more comfortable. Parallel viewing is often more comfortable once it clicks, but it has a firm size limit because your eyes cannot rotate further apart than straight ahead.
Why can I not see the side-by-side effect on my monitor?
Almost certainly because the image is drawn too large, and this catches nearly everyone with parallel viewing. Fusing a parallel pair requires your eyes not to rotate further apart than straight ahead, since human eyes cannot diverge, which caps the separation of matching points at about 63 mm — the distance between your pupils. Each half of a full-width pair on a large monitor is tens of centimetres across, so it cannot be fused at any resolution. The size control under the preview exists for this: shrink it until each half is around 63 mm, or press the fit button. Cross-eyed viewing avoids that divergence limit, though a moderate size is still more comfortable.
Which photos convert best?
Ones with a clear foreground, a distinct middle, and a real background — a landscape with layered distance, a person against a room, a single object on a plain surface. Those give the model unambiguous cues and leave little to invent behind edges. What converts badly is anything flat or ambiguous: a photograph of a photograph, a wall of text, heavy motion blur, large reflections, glass, and shots where subject and background are the same brightness and texture. Very tight close-ups also have little depth structure to work with, so the effect is subtle whatever the strength.
What is the halo or smearing around objects?
Two separate things. Shifting pixels sideways to build a second viewpoint exposes areas that were hidden behind nearer objects, and nothing ever photographed them, so they have to be invented — each gap is patched from the background side of the edge and mirrored inward, but a patch is still a patch, and it grows with strength. Separately, the model works at a much lower resolution than your photo, so its depth edges arrive soft and slightly off the object, which puts a band of background at the wrong depth around heads and shoulders. The edge alignment setting is aimed squarely at the second one.
Does it work on a phone?
Yes, and a phone is the best screen for looking at the result: it is the one display where a parallel side-by-side pair is comfortably inside the width limit, so free-viewing is much easier there than on a desktop monitor. Generating depth is slower on a phone than on a desktop with WebGPU, but it is one photo rather than thousands of frames, so it stays a matter of seconds. Keep the output width modest on mobile — the memory estimate shown before you generate will tell you when the settings are asking for too much.
Related tools
More local, browser-only tools for working with images.