LoRA taught GPT-3 new jobs by training 4.7 million numbers
LoRA taught GPT-3 new jobs by training 4.7 million numbers and leaving 175 billion alone.
How it’s made · 3
Hu et al. 2021, LoRA · Table 4, page 8 · section 4.2, page 5
How do you teach a giant AI model a new job, without building a second giant? In 2021, a team at Microsoft published an answer called Lora. Last lesson, GPT-3, OpenAI's 2020 model, copied answers people wrote by hand. That copying step was fine-tuning. Fine-tuning means more training on a narrow set of examples, so a general model gets good at one job. A dial is one of the numbers inside a model that training turns. The ordinary way nudges all hundred and seventy-five billion of GPT-3's dials. Each tuned job then gets saved as a whole new copy, three hundred and fifty gigabytes. Lora, short for low-rank adaptation, leaves every one of those dials alone. Lora trains four point seven million new numbers instead, and those numbers kept up with the ordinary way.
Every job shares one printed sheet and gets its own clear sheet
Every job shares one printed sheet and gets a clear sheet of its own.
Hu et al. 2021 · Figure 1, page 1 · section 4.1, page 4
Picture GPT-3's dials printed on a huge sheet of paper, one number in every square. Ordinary fine-tuning reprints the whole sheet for every job. Lora locks the printed sheet and lays a clear sheet over the top. Training writes only on the clear sheet. The model reads the two sheets together, so every square counts as the printed number plus whatever the clear sheet adds there.
One GPT-3 grid holds about 151 million dials
One GPT-3 grid holds about 151 million dials.
Brown et al. 2020, Table 2.1 · Hu et al. 2021, section 4.2, page 5 · 150,994,944 dials
GPT-3 keeps its dials in grids. The grids Lora works on are twelve thousand, two hundred and eighty-eight squares wide, and the same number of squares tall. That makes about a hundred and fifty-one million dials in one grid. GPT-3 stacks ninety-six layers, and each layer holds several grids of dials.
A clear sheet written square by square costs as much as the grid
A clear sheet written square by square costs as much as the grid it covers.
LoRA needs a cheaper way to fill the squares.
Hu et al. 2021 · section 4.1, page 4
Lay one clear sheet over one of those grids. If training wrote the clear sheet one square at a time, the clear sheet would need about a hundred and fifty-one million numbers of its own. Do that for every grid in every layer, and the clear sheets weigh as much as GPT-3. The saving has to come from how Lora writes on the clear sheet.
One column times one row fills every square on the clear sheet
One column times one row fills every square on the clear sheet.
Hu et al. 2021 · Figure 1, page 1 · section 4.1, page 4 · 24,576 numbers per pair
Lora writes the clear sheet like a times table. One column of numbers runs down the left edge. One row of numbers runs across the top. Every square gets its column number times its row number. On a small sheet, four squares by four, a column of four numbers and a row of four numbers fill all sixteen squares. On a GPT-3 grid, one column and one row hold twenty-four thousand, five hundred and seventy-six numbers, and those numbers fill all hundred and fifty-one million squares. The numbers you store grow with the edge of the sheet, and the squares they fill grow with its area. Every row on a times-table sheet is the same row, scaled up or down. So Lora can stack a few column-and-row pairs and add their sheets together. The count of pairs is called the rank, and a small count is the low rank in low-rank adaptation.
Only the column and the row learn
Only the column and the row learn, and the 175 billion printed dials sit still.
per layer: 2 grids × 1 pair, or 1 grid × 2 pairs
Hu et al. 2021 · Figure 1, page 1 · section 4.1, page 4 · Table 4, page 8 · Appendix D · 4,718,592 numbers
Training starts with the row full of random numbers and the column full of zeros. Zero times any number is zero, so every square on the clear sheet starts blank. Before training, the model behaves exactly like the original GPT-3. Training then nudges only the numbers in the column and the row. In each of GPT-3's ninety-six layers, the authors used one pair on each of two grids, or two pairs on one grid. Either way, the total came to four point seven million trained numbers. The other hundred and seventy-five billion dials stayed locked.
On GPT-3, LoRA kept up with ordinary fine-tuning in all three tests
On GPT-3, LoRA kept up with ordinary fine-tuning in all three tests.
A tie on the first test, and ahead on the other two.
Hu et al. 2021 · Table 4, page 8 · section 4.2, page 5 · section 4.1, page 4
The authors tested both ways on three jobs. The first job turns a question into a lookup in a database. Lora got seventy-three point four percent of those right, and ordinary fine-tuning got seventy-three point eight. The authors say their scores on that job wobble by about half a percent, so the first job is a tie. The second job judges whether one sentence follows from another. Lora got ninety-one point seven percent right, against eighty-nine point five. The third job sums up a chat, scored by how many words the summary shares with one a person wrote. Lora scored fifty-three point eight, against fifty-two. A bigger Lora setup, with four pairs on the same grids, shrank the saved file for each job from three hundred and fifty gigabytes to thirty-five megabytes. After training, Lora presses each clear sheet onto its printed grid, so the finished model answers as fast as an ordinary fine-tuned one.
Two tests needed only a simple change
On the two tests where the authors varied the pairs, a simple change was enough.
One pair did about as well as 64 pairs.
Hu et al. 2021 · Table 6 and footnote 6, page 10
This picture is roughly right, and here is where the picture breaks. A times-table sheet can only write a simple change. The authors tried from one pair up to sixty-four pairs on the same two grids. On the database job, one pair got seventy-three point four percent right, and sixty-four pairs got seventy-three point five. The authors also say where they expect a few pairs to stop being enough. Their example is a job in a different language from the one GPT-3 learned in. For a job like that, the authors write that retraining the whole model could certainly do better than a thin sheet.
Each 35 MB clear sheet runs on one shared 350 GB GPT-3
Each 35 MB clear sheet runs on one shared 350 GB GPT-3.
Hu et al. 2021 · footnote 4, page 5
Second, a clear sheet only works on top of the printed one. The thirty-five megabyte file holds only the columns and rows. To use that file, you still load the full three hundred and fifty gigabyte GPT-3 underneath. Lora shrinks every extra job instead. For a hundred jobs, a hundred full copies of GPT-3 take about thirty-five terabytes. That is thirty-five thousand gigabytes. One shared GPT-3 plus a hundred clear sheets take about three hundred and fifty-four gigabytes.
LoRA locks the giant’s dials and trains one thin column and one thin row
LoRA locks the giant’s dials and trains one thin column and one thin row per sheet.
4.7 million numbers trained · 175 billion locked
How it’s made · 3
Hu et al. 2021 · Figure 1, page 1
First, lock GPT-3's dials. Lay a clear sheet over some of its grids. Write each clear sheet as a times table, from one column and one row, with the column starting at zero. Train only the columns and the rows. When training ends, press each sheet onto its grid. That is how Lora taught GPT-3 new jobs with four point seven million trained numbers.

















