011 min
Self-Supervised Learning at a glance
- What it is: training a model on unlabeled data by predicting a part that was deliberately hidden.
- Why it exists: most data on the internet has no human-written label, and self-supervised learning turns that unlabeled majority into a training signal.
- What it costs: pretraining BERT-Large took four days on 64 TPU chips, before any task-specific fine-tuning began.
- When it breaks: a pretext task a model can solve by shortcut, instead of by learning real structure, teaches nothing useful.
- As of 2026: self-supervised pretraining is the first step behind nearly every large language model and most modern computer-vision systems.
021 min
The problem Self-Supervised Learning solves
Before self-supervised learning became standard, the strongest models were trained with supervised learning: a person labels each example by hand — this email is spam, this photo shows a dog — and the model learns to predict that label from the input. Supervised learning works, but it needs a labeled example for every pattern the model is meant to learn, and a person has to make each one.
Yann LeCun, a researcher who pushed self-supervised learning into wide use, described the resulting limit plainly: supervised learning is a bottleneck for building more intelligent generalist models that can do multiple tasks and acquire new skills without massive amounts of labeled data. Most of the text, images, and audio that exists carries no label at all — nobody tagged the sentences on an ordinary web page, or captioned every photo ever taken. Supervised learning could only ever train on the small labeled slice of that data.
Earlier attempts to work around the shortage did not remove the limit, only manage it: reuse a small labeled dataset across several related tasks, or spend a labeling budget as carefully as possible. Both still capped what a model could learn at however much labeling a team could afford. Self-supervised learning removes that cap by finding a way to train on data nobody ever labeled, which is most of the data that exists.
032 min
How Self-Supervised Learning works
Self-supervised learning gets its training signal from the data itself, not from a human-written label. The model is given a “pretext task”: part of an example is hidden, and the model has to predict the hidden part from what remains. Because the right answer is just the data before it was hidden, nobody has to label anything — the data supervises itself.
Follow one sentence through the process: “The engineer restarted the server after the deploy failed.” To build a pretext task from it, one word is hidden — say “server” is replaced with a blank. The model sees “The engineer restarted the [blank] after the deploy failed” and has to predict that the missing word is “server.” Doing this well across millions of sentences forces the model to learn grammar, which words tend to follow others, and what kind of thing gets restarted after a failed deploy — none of it written down as a label by a person.
Once this pretraining is finished, the model holds representations — internal patterns that capture structure in the data — that carry over to other tasks. A separate, usually much smaller step called fine-tuning then adapts those representations to a task that does have labels, such as sorting support tickets, using far fewer labeled examples than training a model from nothing would need.
Pretext tasks split into two broad families. A generative one, like the masked-word task above, asks the model to reconstruct the hidden part directly. A contrastive one instead asks the model to tell which pairs of examples are different versions of the same underlying thing and which are not, without reconstructing anything. Both families can be paired with the same kind of underlying architecture — the Transformer, for instance, is trained with a generative pretext task in one well-known model and can be paired with other pretext tasks elsewhere — so the pretext task and the architecture are separate design choices.
BERT, one of the best-documented examples of this approach, actually runs two pretext tasks together. Alongside predicting masked words, it is given two sentences, A and B, and has to decide whether B genuinely follows A in the source text or was pulled at random from elsewhere in the corpus. Specifically, when the training example is built, 50% of the time B is the actual next sentence that follows A (labeled as IsNext), and 50% of the time it is a random sentence from the corpus (labeled as NotNext). This second task, called next-sentence prediction, is meant to teach the model something about how sentences relate to each other, on top of what the masked-word task teaches about words inside one sentence.
041 min
A concrete example: BERT
BERT, published by Jacob Devlin and colleagues at Google in 2018, applies this idea directly to text. During pretraining, roughly 15% of the words in each sentence are hidden, and the model has to predict each hidden word using the words on both sides of it — a pretext task called masked language modeling — alongside the next-sentence-prediction task described above.
Pretraining ran on unlabeled text: articles from the web and the text of digitized books, with nobody labeling any of it. It obtains new state-of-the-art results on eleven natural language processing tasks, including pushing the GLUE score to 80.5% (7.7% point absolute improvement) over the previous best result. The next-sentence-prediction task alone reached 97%-98% accuracy once trained, a sign of how much structure a model can pull out of a pretext task that needed no person to write a single label.
Training BERT’s larger model, BERT-Large, was performed on 16 Cloud TPUs (64 TPU chips total), and each pretraining run took 4 days to complete, before any of the eleven fine-tuning runs started. That four-day run only had to happen once: the same pretrained model was then fine-tuned separately, cheaply, for each of the eleven tasks, which is the entire economic case for paying the large cost up front.
052 min
A second case: SimCLR and images
Two years later, Ting Chen and colleagues at Google showed the same underlying move works in a completely different kind of data: photographs, where there is no obvious missing word to predict. SimCLR’s pretext task starts from one unlabeled image, makes two different distorted copies of it — one cropped and recolored, one flipped and blurred — and trains the model to recognize that both distorted copies came from the same original photo, picked out of a large batch of otherwise unrelated images.
A linear classifier trained on self-supervised representations learned by SimCLR achieves 76.5% top-1 accuracy, which is a 7% relative improvement over previous state-of-the-art, matching the performance of a supervised ResNet-50 trained on labels from the start. When fine-tuned on only 1% of the labels, the same representations achieve 85.8% top-5 accuracy, outperforming AlexNet, a well-known model trained with 100 times more labeled data.
The pretext task changed completely between BERT and SimCLR — predicting a missing word against recognizing two crops of one photo — but the underlying move, turning the data’s own structure into the training signal, stayed the same. The choice of distortion in SimCLR’s pretext task matters more than it looks. Cropping and recoloring an image forces the model to recognize the same object or scene despite a changed viewpoint or lighting, which is close to the actual variation a useful visual representation needs to handle. A weaker distortion — one that left an obvious shortcut, such as a fixed border added by the cropping tool — would let the model solve the pretext task without learning anything transferable.
062 min
What self-supervised pretraining means for your work
For a founder or PM deciding whether a product needs a model trained from scratch, self-supervised pretraining is usually the reason the answer is no. A pretrained model already encodes a large amount of general structure, so the cheaper move is fine-tuning it on a much smaller labeled dataset built for the product, rather than starting from nothing.
For an engineer choosing which base model to build on, the pretext task a model was trained with says a lot about what it will and will not already be good at. A model pretrained mostly on masked-word prediction from web text has picked up broad language structure but nothing about a product’s specific terms, which is exactly the gap fine-tuning or retrieval-augmented generation is meant to close.
For a designer or researcher reading a vendor’s claim that a model “understands” a domain, the honest reading is usually narrower: self-supervised pretraining exposed the model to a large amount of related text or images, which is different from having been taught the specific facts a product needs to get right.
This also changes how a team should spend its own labeling effort. Because a pretrained model already carries broad structure, the labeled examples a team collects are best spent on the gap between that general structure and the specific task — the edge cases and product-specific terms that generic web text or generic photos never covered — rather than re-teaching the model things its pretraining already handled.
071 min
What Self-Supervised Learning costs
Pretraining is the expensive part, and it is paid once. BERT-Large trained on 64 TPU chips for four days before a single fine-tuning run began, a cost measured in dedicated hardware-days rather than the minutes a typical fine-tuning pass takes afterward. Modern large language models scale the same idea to weeks of training across thousands of accelerators, which is why pretraining a foundation model from scratch stays rare outside a small number of large labs.
Fine-tuning a pretrained model is comparatively cheap: a few hours on a single GPU is enough to adapt a BERT-sized model to a new task. That gap between the two costs is the entire economic case for self-supervised pretraining: pay the large cost once, then reuse the result across many tasks.
081 min
What Self-Supervised Learning does not solve
A pretext task that is too easy defeats the point. If the hidden part of an example can be guessed from a shortcut unrelated to the structure the designer wanted the model to learn — a fixed image border left by the cropping tool that reveals which crop came from which photo, for instance — the model learns the shortcut instead, and the resulting representations transfer poorly to real tasks.
Self-supervised pretraining also does not remove the need for labeled data; it reduces how much is needed. A model pretrained this way still needs task-specific fine-tuning, and that step still needs enough labeled examples to reliably teach the task at hand, especially when the task looks little like the pretraining data.
It is not a substitute for choosing the right pretraining data, either. A model pretrained only on general web text or generic photographs transfers well to tasks that resemble that data and poorly to tasks that do not. Self-supervised learning removes the labeling bottleneck; it does not remove the need to think about what the model actually saw before it met a labeled example.
091 min
Self-Supervised Learning vs. nearby concepts
Self-supervised learning is often lumped together with unsupervised learning, and the two share a starting point: no human-provided labels. The difference is in the training signal.
| Supervised learning | Unsupervised learning | Self-supervised learning | |
|---|---|---|---|
| Labels | Human-written, one per example | None | None |
| Training signal | The human label | Structure found directly in the data, such as clusters | A prediction target built from the data itself, such as a hidden word |
| Example task | Classify an email as spam | Group similar customers by behavior | Predict a masked word |
Unsupervised learning finds structure directly, with no explicit prediction target at all. Self-supervised learning manufactures an explicit prediction target out of the data itself — the hidden word, the matching image crop — then trains with the same predict-and-correct process supervised learning uses. It sits between the two: unsupervised in that it needs no human labels, supervised in the mechanics of how it actually learns.
It is also easy to conflate with fine-tuning, but the two are separate stages of one pipeline: self-supervised pretraining builds general representations from unlabeled data first, and fine-tuning adapts those representations to one specific labeled task afterward. A model is self-supervised during pretraining and ordinarily supervised during fine-tuning.
?5 questions
Questions people ask
Is self-supervised learning the same as unsupervised learning?
When should I use self-supervised pretraining instead of training a model from scratch with labels?
What are the risks or limits of self-supervised learning?
Does self-supervised learning require GPUs or TPUs?
Do I still need labeled data if I use self-supervised learning?
§4 sources
Sources
Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2018/2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv:1810.04805.
Chen, T., Kornblith, S., Norouzi, M., & Hinton, G. (2020). A Simple Framework for Contrastive Learning of Visual Representations. arXiv:2002.05709.
LeCun, Y., & Misra, I. (2021). Self-supervised learning: The dark matter of intelligence. Meta AI Blog.
Self-supervised learning. Wikipedia.





