Deep Dive
The Machine Gaze
Chapter 8 of the tour covers this ground in a page: a face goes into a machine-learning system, a claim about identity, character, or orientation comes out, and the claim borrows the authority of statistics and neural networks that physiognomy never had. Two studies get roughly a paragraph each — Wang and Kosinski's 2018 claim to detect sexual orientation from dating-profile photographs, and Kosinski's 2021 follow-up claiming to detect political orientation the same way — before the chapter turns to a study of the opposite kind, Gender Shades, and closes on deepfakes as an unsettling endpoint. This deep-dive slows down on each of those pieces: what Wang and Kosinski's model actually was and what it was trained on, the specific methodological critique that took it apart, the same pattern repeating in Kosinski's later political-orientation study, the real audited accuracy gaps Gender Shades found in commercial gender-classification systems, and what deepfake generation actually is as a technology and as an evidentiary problem.
It also opens further back than Chapter 8 does, with Francis Galton's 1878 composite photographs — covered elsewhere in this tour as an origin point for physiognomy's "criminal type" hunting, not as a prologue to machine learning. Worth being explicit up front: two of the four studies below are debunked pseudoscience dressed in a lab coat, and one is real, credible, currently influential science that says something true about where these systems fail. The deep-dive treats them differently on purpose, because the evidence treats them differently.
The same move, a century and a half apart
In 1878, Francis Galton described composite photography: superimposing multiple exposures of different individuals' faces, aligned on the eyes, onto a single photographic plate to produce one blended, averaged portrait.1 He used it hunting a "criminal type," combining photographs of convicted prisoners sorted by offense and hoping a shared, visible physiognomy would emerge from the blend. It did not — the composites came out smoother and more generic than any individual face that went into them, not more criminal-looking, and Galton's own attempt at that particular project failed on its own terms.
The underlying move Galton was making — extract a statistical pattern from a large set of faces, then claim that pattern reveals something about character, identity, or type that a single face alone would not show — did not die with his failed composites. It is the same move a deep neural network makes when it is trained on thousands of labeled facial images and asked to output a probability that the person pictured is gay, or conservative, or a criminal: aggregate the faces, find whatever regularities separate the labeled groups, and present the output as a discovery about the face itself. Galton did this by hand, with a camera and a stack of glass plates, and got a negative result he was honest about. The systems covered in the rest of this deep-dive do the same thing with far larger datasets and far more computational power — and inherited the same vulnerability: a pattern found in a dataset is not automatically a pattern that lives in bone, muscle, or skin. It can just as easily be a pattern in how the photographs were taken.
The "AI gaydar" study, in more depth
Michal Kosinski and Yilun Wang built their classifier on 35,326 photographs scraped from a US dating website, drawn from self-identified gay and heterosexual users, and used a standard pre-trained deep neural network (VGG-Face) to extract facial features from each image before feeding those features into a simpler classifier. On single images, the system distinguished gay from heterosexual men with 81% accuracy and women with 71% accuracy — far above the roughly 61% and 54% that human judges managed on the same photos — and the paper argued the gap supported a "prenatal hormone theory" under which sexual orientation leaves a physical signature in facial structure.2
The rebuttal that took the study apart came a few months later, from Blaise Agüera y Arcas and Margaret Mitchell — both Google AI researchers at the time — and Alexander Todorov, then a Princeton psychologist who has spent his career studying what faces do and don't communicate. Their published critique walked through the dating-profile dataset image by image and found a consistent, mundane explanation hiding inside the "sophisticated" one: gay and heterosexual users differed systematically in eyeshadow and makeup use, facial hair styling, eyewear, head angle and framing in selfies, and even how much sun-exposed skin tone their photos showed — all self-presentation and camera choices, not facial bone or soft-tissue structure. When they built a model using nothing but a handful of yes/no questions about those surface traits, it matched the deep neural network's accuracy almost exactly.3 A model that can be matched by a checklist of grooming and camera-angle questions was never detecting a biological signature in the first place. The honest reading isn't "faces encode sexual orientation" — it's that a well-groomed, self-selected dataset let a pattern-matching system detect grooming, and the paper's framing mistook that for detecting a person.
Reading politics off a face, and the same collapse
Kosinski returned to the same premise in 2021, this time working alone. Drawing on more than a million photographs — dating-site profiles from the US, UK, and Canada plus Facebook profiles from the US — he ran a face-recognition algorithm over matched pairs of self-identified liberal and conservative faces and reported that it picked the correct political orientation 72% of the time — beating both chance and the human judges tested on the same pairs — and again framed the result as evidence that political ideology has a detectable facial signature.4
Reading political affiliation off a face was not a new idea in 2021. A decade earlier, the psychologists Nicholas Rule and Nalini Ambady had shown people cropped photos of US Senate candidates and found they could guess party at about 57% accuracy — a small but real effect the authors traced to stereotyped impressions of traits like warmth and dominance, not to any claim about facial structure.5 The jump from that modest, human-judge effect to Kosinski's 72% tracks the size and self-selection of his dataset far better than it tracks any new biological signal.
The methodological objection landed in almost identical terms to the one that had already taken apart the sexual-orientation paper, and came in part from the same critic: Alexander Todorov argued publicly that the algorithm's apparent success very plausibly reflected non-facial clues baked into self-posted dating and social-media photographs rather than any signal in facial structure itself — head angle, expression, framing, and the styling and presentation choices that correlate, on average, with political and cultural identity in exactly the way they correlate with countless other social groupings. Kosinski's own attempt to control for this by tightly cropping and standardizing the images did not resolve the objection, because cropping a photo more tightly does not remove a subject's chosen head tilt, expression, or grooming from the crop — it just moves those cues closer to the center of the frame. A face-recognition system finding a statistical regularity in a large set of self-posted photographs is not thereby detecting a stable biological trait. Treating "the model found a pattern" as proof the pattern lives in bone and muscle repeats, point for point, the leap physiognomy made two centuries earlier — laundered, this time, through a neural network instead of a caliper.
A different, well-founded finding: Gender Shades
Not every claim about face-reading machines collapses under scrutiny. In 2018, Joy Buolamwini and Timnit Gebru audited three commercial gender-classification systems, built by IBM, Microsoft, and the Chinese firm Face++, against a benchmark set of faces they built specifically to be balanced across skin tone and gender, using the Fitzpatrick scale to classify skin type rather than relying on the more racially coded, and less evenly distributed, categories used in most existing face datasets. All three systems classified lighter-skinned men with near-perfect accuracy — around 99%. All three did dramatically worse on darker-skinned women: IBM's system misclassified darker-skinned women's gender 34.7% of the time, Face++'s did so 34.5% of the time, and Microsoft's 20.8% of the time, against error rates under 1% for lighter-skinned men in the worst-performing case.6
The gap traced directly to training and benchmark data: the face datasets these systems (and most others in the field at the time) were built and evaluated on skewed heavily toward lighter-skinned faces, so the systems had, in effect, seen far more examples of some faces than others during development. This is established, widely replicated, currently influential science — it reshaped how the field audits facial-analysis systems for bias, prompted several of the audited companies to curtail or overhaul their facial-analysis products, and it says something true and useful that neither Wang and Kosinski's study nor Kosinski's political-orientation study says: that a face-reading system's errors are not random noise but a direct reflection of whose faces were, and weren't, well represented in the data that trained it. Gender Shades doesn't claim a face reveals hidden character or identity. It measures, carefully and reproducibly, where a specific set of machines fails — which is exactly the kind of claim the debunked studies above never made and could not have supported.
The accuracy gap was not just a benchmark number. Later that year, the ACLU ran a different commercial system — Amazon's Rekognition — against a database of 25,000 arrest photos and found it falsely matched 28 sitting members of Congress to mugshots, with the false hits landing disproportionately on members of color (nearly 40% of the errors, against 20% of Congress).7 Different tool, different task, same lesson about whose faces these systems handle worst.
The face as evidence, in an age of synthetic faces
The same broad family of technology that misreads faces can now manufacture them. The technique underneath most modern deepfakes traces to a 2014 paper by Ian Goodfellow and colleagues describing generative adversarial networks: two neural networks trained against each other, one generating images and the other trying to tell generated images from real ones, improving in tandem until the generator produces images the discriminator can no longer reliably flag as fake.8 Applied to faces, that idea and its descendants can now generate photorealistic images and video of people saying and doing things that never happened, at a fidelity that routinely defeats casual human judgment and increasingly strains automated detection tools built to catch it.
The word itself came not from a lab but from a screen name. In December 2017 the journalist Samantha Cole tracked down the anonymous Reddit user "deepfakes," a self-described hobbyist programmer who had used open-source machine-learning tools to swap celebrities' faces into pornographic videos — the first public application of the technique, and the source of the name that stuck.9
Legal scholars Robert Chesney and Danielle Citron mapped the resulting problem in detail in 2019: once synthetic video is convincing and cheap to produce, the evidentiary value of "there's a video of it" starts to erode in both directions. Fabricated video can frame someone for something they never did, and — the effect the authors named the "liar's dividend" — the mere existence of convincing fakes gives genuinely guilty people a plausible way to dismiss real, incriminating footage of them as fabricated.10 Put next to the rest of this deep-dive, that is the uncomfortable endpoint of the machine gaze: a century and a half after Galton tried to average a criminal type out of a stack of photographs, the tools built to read faces automatically are being outpaced by tools that can fabricate faces automatically — a face, however confidently a person or a system reads it, was never a transparent window onto identity, character, orientation, or truth, and now it isn't even reliable proof of who was actually in front of the camera.
The recurring lesson
Lay this deep-dive's chain next to the rest of the tour and the pattern is the same one, running the whole length of it. Lavater read moral character off a jawline. Camper and Retzius converted that impression into a bony angle and a cranial ratio. Broca built a measuring program that made the numbers look rigorous without making the underlying premise any truer. Morton and Lombroso used that apparent rigor to rank populations and invent a criminal "type." Boas, a generation later, took actual measurements of thousands of real immigrant families and showed head shape shifting within a single generation — real data, gathered honestly, doing the opposite work: dismantling a fixed-type claim instead of manufacturing one. Wang and Kosinski, and Kosinski again in 2021, repeated the original mistake with a neural network standing in for the caliper, and Galton's own composite photographs, a century and a half earlier, had already previewed both halves of that repetition — a real technique, aimed at a target that was never there to find.
The corrective, each time it has actually worked, has looked the same, too: not a better instrument for reading faces, but a check on whether the instrument's output reflects the thing it claims to measure or something else riding along with it — Lichtenberg's plain-language rebuttal of Lavater in 1778, Boas's re-measurement of Morton's premise a century later, and Agüera y Arcas, Mitchell, and Todorov's demonstration that a checklist of grooming questions could match a "sophisticated" neural network. Gender Shades belongs on that same corrective side of the ledger, not the debunked side — it is what rigorous, bias-audited measurement of a face-reading system actually looks like, applied to real commercial products with real stakes. That is the throughline this whole tour has been tracing: a face, read quickly and confidently, keeps seeming to reveal an inner truth about the person wearing it. It never has. What has repeatedly worked, from Lichtenberg to Boas to the researchers who took apart the AI-gaydar study, is the slower, harder discipline of checking the reading against the actual evidence — and being willing to conclude, again, that the face was never saying what it appeared to say.
Further reading
The Machine Gaze
The shorter tour version of this history: AI physiognomy, the old mistake, automated.
Reading the Face
The history of physiognomy and craniometry, and the harm it did — the tradition this deep dive's algorithms unknowingly repeat.
The Skull That Broke Race Science
Boas 1912 and the re-analysis war that still isn't settled.
References
- (1878). Composite portraits, made by combining those of many different persons into a single figure. Nature, 18(447), 97–100. ↩
- (2018). Deep Neural Networks Are More Accurate Than Humans at Detecting Sexual Orientation From Facial Images. Journal of Personality and Social Psychology, 114(2), 246–257. ↩
- (2018). Do Algorithms Reveal Sexual Orientation or Just Expose Our Stereotypes?. Medium. ↩
- (2021). Facial Recognition Technology Can Expose Political Orientation From Naturalistic Facial Images. Scientific Reports, 11, Article 100. ↩
- (2010). Democrats and Republicans can be differentiated from their faces. PLoS ONE, 5(1), e8733. ↩
- (2018). Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification. Proceedings of Machine Learning Research, 81, 77–91. ↩
- (2018). Amazon’s Face Recognition Falsely Matched 28 Members of Congress With Mugshots. American Civil Liberties Union, July 26, 2018. ↩
- (2014). Generative Adversarial Networks. Advances in Neural Information Processing Systems, 27, 2672–2680. ↩
- (2017). AI-assisted fake porn is here […]. Motherboard / VICE, December 11, 2017, tracing the term “deepfake” to an anonymous Reddit user. ↩
- (2019). Deep Fakes: A Looming Challenge for Privacy, Democracy, and National Security. California Law Review, 107(6), 1753–1819. ↩