Cameras and the Autocomplete
Unable to bend the laws of optics, modern smartphone cameras trade optical reality for math to function as statistical autocomplete engines.
By Hugh Askell
Sparked by Eclipse: The Xiaomi 17 Ultra Confuses the Moon and the Sun · discussion

There was an amusing piece in Frandroid recently about a bug in the Xiaomi 17 Ultra. A user pointed their phone at a solar eclipse, tapped the shutter, and the device helpfully output a high-res, perfectly textured full moon. Oops. The software detected a bright circle in a dark sky and, as the publication pointed out, the algorithm simply applies its moon filter the moment it detects a luminous circle against a black background — pasting an entirely fabricated celestial body over a rare astronomical event. This was predictably mocked on Hacker News as a silly software glitch, but that sort of tech-press schadenfreude usually misses the underlying physics. We are looking at the exact moment a structural limit is exposed.
To understand why a thousand-dollar computer thinks the sun is the moon, you have to bypass the software layer entirely and look at the physical supply chain. Consumers stubbornly refuse to buy a phone thicker than about 8 millimetres, which creates a rather brutal bottleneck for the optical engineers. In traditional photography, capturing distant detail requires a long focal length, meaning physical distance between the glass lens and the sensor. You cannot simply bend the immutable laws of glass optics to fit inside a flat metal slab. If you want a genuine 100x optical zoom that captures true photons from a distant object, you need a camera bump the size of a coffee mug. We saw early, awkward attempts at bridging this gap a decade ago, like the Samsung Galaxy S4 Zoom, which essentially bolted a smartphone to a compact digital camera. Nobody bought it.
Because the handset makers cannot change the physics of light, and because the industry still has to deliver a compelling marketing upgrade every twelve months to justify a massive global hardware cycle, the smartphone sector hit a wall. You can only buy so much raw performance from sensor manufacturers like Sony. The only way forward, therefore, was to dismantle the traditional optical pipeline entirely and substitute raw compute.
For the last half-decade, the device in your pocket has operated essentially as a statistical prediction engine attached to a piece of glass. It captures a noisy, physically constrained input, cross-references that shape against an on-device model, and queries a database to fill in the missing pixels. We have seen this exact structural workaround before, most famously with the Samsung Space Zoom scandal of 2023, where users realised their impressive lunar photography was basically being painted onto the viewfinder in real-time. When you reach the absolute Z-axis limit of the hardware, you abstract the problem away into software, turning an optics problem into a database problem. You trade photons for math.
Naturally, the tech industry likes to obscure this entire process behind the quasi-religious jargon of artificial intelligence. We get breathless marketing copy about neural engines and computational mastery. When the system makes an error like the Xiaomi eclipse bug, commentators say the network is hallucinating or making a cognitive mistake, anthropomorphising a piece of consumer electronics. This is mostly nonsense. A neural network lacks any intrinsic understanding of a moon or an eclipse, much like a household appliance has no concept of dirty socks.
This software is a statistical sorting mechanism that has been fed millions of images of high-contrast circles in dark skies and instructed to apply a probabilistic, high-resolution texture to a low-resolution geometric match. You have accidentally constructed a circle-texturer rather than a camera, and when you pointed it at a dark circle surrounded by a corona, it ran the exact mathematical routine it was programmed to execute.
Stepping back to a thirty-year or even a 150-year industry curve, this represents a profound epistemological shift. Photography has always been understood as an objective physical record — photons hitting silver halide, or later, photons hitting a CMOS sensor and being converted directly into electrical values. (Obviously there were always hoaxes, like the Cottingley Fairies, or the elaborate airbrushing of the Soviet era, but they were anomalies that required deliberate darkroom manipulation, rather than the default operating state of the equipment.) A camera was a device that recorded the light that was actually there.
We have now structurally abandoned that premise. To bypass the immovable laws of glass, the smartphone camera has become a statistical autocomplete engine. It guesses what the scene should look like based on prior training data, heavily augmenting the sparse photons that actually made it through the 8mm lens. Every time you press the shutter button, the processor evaluates the light, decides what category of scene you are probably looking at, and paints in the details that the physical hardware was simply too small to capture. In effect, we traded optical reality for probabilistic aesthetics.
So, where does the boundary sit? If the device smooths out a noisy shadow, we accept it as photography. If it sharpens an edge, we call it image processing. What if it replaces a blurry eclipse with a stock texture of a lunar crater? The question isn't whether the software will eventually stop confusing the sun for the moon—it will. But if a photograph is now just a low-resolution prompt for a probabilistic database, do we still agree on what it means to take a picture? I don’t think we even agree on the vocabulary.