Phone cameras face severe physical constraints. Sensors are small, lenses are tiny, and there's almost no space for optics.

By any traditional optical reasoning, images from such a system should be substantially worse than they are. The gap is filled by computation.

The physical constraint

Image quality depends substantially on how much light reaches the sensor, which depends on sensor area and aperture.

A phone sensor is a fraction of the size of one in a dedicated camera. Less light means more noise, less dynamic range and worse low-light performance.

Depth of field is also determined by physics. A small sensor with a short focal length produces images where nearly everything is in focus, which is why phone photos naturally lack the background separation that larger cameras produce optically.

These constraints haven't gone away. They've been worked around.

Multi-frame processing

The central technique and the reason modern phone images look the way they do.

When you press the shutter, the phone doesn't capture one image. It has been continuously capturing frames, and it selects and combines several.

Combining multiple exposures reduces noise, because noise is random and averages out while signal doesn't. This is why phone images in moderate light are cleaner than the sensor size should permit.

Frames at different exposures are combined to extend dynamic range, capturing both bright sky and shadow detail in a single output.

And alignment between frames compensates for hand movement, allowing effectively longer exposures than would otherwise be possible.

Night modes are the visible extreme of this — capturing many frames over several seconds and combining them.

The processing pipeline

Beyond frame combination, extensive processing is applied.

Noise reduction, which trades detail for cleanliness. Sharpening, which trades naturalness for apparent detail. Tone mapping, which compresses a wide dynamic range into a displayable image. Colour adjustment towards what's expected rather than what was measured. Local adjustments to specific regions — faces, sky, foliage — identified by scene analysis.

Each of these is a decision, and the aggregate is why images from different manufacturers look distinctly different from the same scene.

It's also why the characteristic phone look exists: high local contrast, saturated colour, everything sharp, minimal shadow. That's a set of processing choices calibrated to look good on a small bright screen.

Portrait mode

Worth explaining because it's the clearest example of computation substituting for optics.

Background blur in a traditional camera is a physical consequence of a large aperture and a large sensor. Phones cannot produce it optically.

Instead, the phone estimates a depth map — using multiple lenses, dedicated depth sensors, or machine learning from a single image — segments the subject, and applies a synthetic blur.

The characteristic failures follow from this. Hair, glasses, gaps between limbs and complex edges are where segmentation errs, producing the sharp cutouts and inconsistent blur that give it away.

The blur itself is also synthetic, which is why it lacks the optical characteristics real defocus produces.

Zoom

Another area where marketing and physics diverge.

Optical zoom requires physically different focal lengths, which in phones means separate cameras. A phone with a telephoto module has genuine optical reach at that focal length only.

Everything between the fixed focal lengths is cropping and upscaling — digital zoom, however it's labelled.

High zoom figures generally represent heavy computational upscaling, which can produce impressive-looking results and involves synthesising detail that wasn't captured. What you see at extreme zoom is partly generated rather than recorded.

What this means for photographers

Practical consequences.

Shooting raw bypasses most of it. Raw files contain sensor data before processing, giving control over decisions the phone makes automatically. The result initially looks worse — flat, noisy — because you're seeing what the sensor captured rather than what the processing produced.

Some phones offer a computational raw format, retaining multi-frame benefits while allowing later adjustment, which is generally the better option.

The processing assumes typical scenes. Unusual lighting, high-contrast subjects and intentionally dark images fight the processing, which will try to correct what it interprets as a mistake.

Motion is the remaining weakness. Multi-frame techniques assume the scene is broadly static between frames. Moving subjects produce artefacts, which is why phone photos of children and animals in low light remain unreliable.

Light still matters most. No processing substitutes for good light. The single largest improvement available to any phone photographer is paying attention to where the light is coming from.

Where it's going

The trajectory is towards more generation and less capture. Features that reconstruct detail, remove objects, adjust expressions and extend beyond the frame are already deployed.

Which raises a question that's more than technical: at what point does an image stop being a record of what was there.

Several manufacturers have begun embedding provenance metadata indicating what processing was applied, and standards work in this area is ongoing. It's likely to matter more as the gap between captured and generated content narrows.