← Evolution of the Face

Chapter 08 of 08 · 7 min read

The Machine Gaze

Machine vision does not escape physiognomy merely by replacing calipers with neural networks.

Chapter 8

The Machine Gaze

An old claim in new clothes

This tour opened with physiognomists reading character off a jawline and craniometrists reading intelligence off a skull's cubic capacity, and it passed through a chapter on just how little the confident, instant impression a face makes actually tracks the truth of the person wearing it. The same pattern shows up again in the 2010s, dressed in a lab coat. A face goes into a machine-learning system; a claim about identity, character, or orientation comes out; the claim is presented with the borrowed authority of statistics and neural networks. The stakes are no longer phrenology cabinets and skull calipers, but the underlying move is identical — and it fails in the same way, for the same reason: the system was never measuring what it claimed to measure. The most literal revival came in 2016, when Xiaolin Wu and Xi Zhang posted a paper claiming a classifier could tell convicted criminals from non-criminals by their ID photos alone — Lombroso's "born criminal" (Chapter 2) restated almost word for word, with a neural network standing in for the calipers.1 It drew a swift, well-documented backlash from AI researchers, who noted among other flaws that the "criminal" and "non-criminal" photos differed in mundane ways a classifier could latch onto, down to the non-criminals being more likely to be smiling.2

The "AI gaydar" claim, and its collapse

In 2017 and 2018, the Stanford researchers Yilun Wang and Michal Kosinski published a study claiming that a deep neural network trained on dating-profile photographs could distinguish gay from heterosexual men and women more accurately than human judges could — 81% accuracy for men, 71% for women, against much lower human rates — and framed the result as support for a theory that sexual orientation leaves a detectable signature in facial structure.3 The claim did not hold up. Critics — including Google AI researchers and social psychologists studying face perception — pointed out that the classifier had access to grooming choices, makeup, facial hair style, camera angle, lighting, and even photo filters that differ systematically between the dating-profile photos gay and straight people choose to post, and that these self-presentation and image-quality confounds, not any biological signal in bone or soft-tissue structure, were sufficient to explain the model's accuracy. When those confounds were controlled for or stripped out, the system's advantage over chance collapsed.4 The honest reading of the result isn't "faces encode sexual orientation" — it's that a well-groomed dataset let a pattern-matching system detect grooming and presentation, and the paper's framing mistook that for detecting a person.

Reading politics off a face

Kosinski returned to the same premise in 2021, publishing a study claiming that a facial recognition algorithm could predict political orientation from photographs with 72% accuracy across more than a million images drawn from the US, UK, and Canada.5 It drew the identical methodological critique as the sexual-orientation paper, for the identical reason: politically liberal and conservative people differ, on average, in how they present themselves for a camera — head angle, facial expression, glasses, styling choices that correlate with political and cultural identity — and a large enough training set lets a classifier pick up on those presentation differences without ever touching an innate "political face." That is not the end of the story, though: a 2024 follow-up by Kosinski's own team used standardized photographs with expression, head angle, and self-presentation controlled and still found a small but above-chance association — so whether the political-face signal is entirely a presentation artifact or reflects something more durable remains an open, contested question.6 Facial recognition systems are very good at finding statistical regularities in large image sets; they are not thereby detecting a stable biological trait, and treating "the model found a pattern" as proof that the pattern lives in bone and muscle repeats exactly the leap physiognomy made two centuries earlier, just laundered through a neural network instead of a caliper.

A different, well-founded finding: Gender Shades

Not every claim about face-reading machines collapses under scrutiny — some of the most important findings in this space are the opposite of overclaiming: they show where these systems reliably fail. The audit that established this began with a mask: as an MIT graduate student, Joy Buolamwini found a face-analysis system would not detect her dark-skinned face until she held a plain white mask up to it — an experience she credits as the spark for the work that followed.7 In 2018, Buolamwini and Timnit Gebru audited three commercial gender-classification systems and found their error rates were not evenly distributed. Accuracy was highest for lighter-skinned men and dropped sharply for darker-skinned women, with error rates for that group reaching up to 34.7% versus under 1% for lighter-skinned men in the worst-performing system — a gap traced to training datasets that skewed heavily toward lighter-skinned faces.8 This is established, widely replicated, currently influential science — it reshaped how the field audits facial-analysis systems for bias, and it says something true and useful: that these systems' errors are not random noise but a direct reflection of whose faces were, and weren't, well represented in the data that trained them.

A plainer version of the same failure had already made the news. In 2015 Google Photos' new auto-tagging labeled photographs of two Black users as "gorillas." Google's fix was not to retrain the model on more representative faces but to delete "gorilla," "chimp," and "monkey" from the app's vocabulary — a patch still in place more than two years later.9

The face as evidence, in an age of synthetic faces

The accuracy gap Buolamwini and Gebru measured has a human cost with a name. In January 2020 Detroit police arrested Robert Williams at his home, in front of his family, after a facial-recognition search matched a blurry shoplifting still to his driver's-license photo; he spent thirty hours in a cell before anyone checked the match against his actual face — the first publicly documented wrongful arrest in the US traced to a face-recognition error.10

The same technology that misreads faces can now manufacture them. Deepfake systems generate photorealistic faces of people saying and doing things that never happened, at a fidelity that routinely defeats casual human judgment and increasingly strains automated detection tools as well.

The technique behind most modern deepfakes has a famously offhand origin. Ian Goodfellow has said the core idea — pitting two neural networks against each other — came to him in 2014 during a bar argument with friends over how to make computers generate realistic images; he went home and coded a working version that night.11

Put next to the rest of this chapter, that's the uncomfortable endpoint of the machine gaze: a century and a half after physiognomists claimed a face could reveal a person's inner truth, the tools built to read faces automatically are being outpaced by tools that can fabricate faces automatically — which makes the underlying lesson of this whole tour more urgent, not less. A face, however confidently a person or a system reads it, is not a transparent window onto identity, character, orientation, or truth. It never was.

Further reading

References

  1. Xiaolin Wu & Xi Zhang (2016). “Automated Inference on Criminality Using Face Images.” arXiv:1611.04135.
  2. Blaise Agüera y Arcas, Margaret Mitchell & Alexander Todorov (2017). Physiognomy’s New Clothes. Medium, May 6, 2017.
  3. Wang, Y., & Kosinski, M. (2018). Deep Neural Networks Are More Accurate Than Humans at Detecting Sexual Orientation From Facial Images. Journal of Personality and Social Psychology, 114(2), 246–257.
  4. Blaise Agüera y Arcas, Alexander Todorov, & Margaret Mitchell (2018). Do Algorithms Reveal Sexual Orientation or Just Expose Our Stereotypes?. Medium.
  5. Kosinski, M. (2021). Facial Recognition Technology Can Expose Political Orientation From Naturalistic Facial Images. Scientific Reports, 11, Article 100.
  6. Michal Kosinski, Poruz Khambatta & Yilun Wang (2024). Facial recognition technology and human raters can predict political orientation from images of expressionless faces even when controlling for demographics and self-presentation. American Psychologist, 79(7), 942–955.
  7. Joy Buolamwini (2023). Unmasking AI: My Mission to Protect What Is Human in a World of Machines. New York: Random House.
  8. Buolamwini, J., & Gebru, T. (2018). Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification. Proceedings of Machine Learning Research, 81, 77–91.
  9. Jackie Snow (2018). Google Photos still has a problem with gorillas. MIT Technology Review, Jan 11, 2018. Original 2015 incident reported by Jacky Alciné.
  10. Tate Ryan-Mosley (2021). The new lawsuit that shows facial recognition is officially a civil rights issue. MIT Technology Review, Apr 14, 2021; ACLU, Williams v. City of Detroit.
  11. Martin Giles (2018). The GANfather: The man who’s given machines the gift of imagination. MIT Technology Review, Feb 21, 2018.