EconReads
Donate

The Economics of Artificial Intelligence

Data: The Raw Material of AI

Why AI models need vast amounts of data, the legal fights over using copyrighted work for training, and concerns about running out of high-quality data.

AI models learn from data. Large language models are trained on enormous amounts of text from books, websites, articles and code. Image models learn from huge collections of pictures. Data has become a key economic input.

Where the data comes from

Much training data has been gathered from the public internet, through large collections of web pages. Developers also use licensed data, public domain works, and data created by humans specifically for training, such as people rating AI answers.

Many authors, artists, news organisations and publishers argue that training AI on their work without permission or payment violates copyright.

  • In December 2023, The New York Times sued OpenAI and Microsoft, alleging they used millions of its articles to train AI models without permission.
  • Authors and artists have filed lawsuits against several AI companies.
  • In 2025, AI company Anthropic agreed to pay 1.5 billion dollars to settle a class action lawsuit by authors over the use of pirated books in training, one of the largest copyright settlements ever.

AI companies argue that training is fair use, since models learn patterns rather than copying works, and that restricting it would slow innovation. Courts in different countries are still deciding these questions.

Licensing deals

Some AI companies have signed deals to license content from news publishers, image libraries and online forums, paying for access. This is creating a new market for data.

Running out of data?

Researchers at the organisation Epoch AI have estimated that AI developers could use up much of the high-quality public text on the internet within the coming years at current growth rates. This has led companies to explore synthetic data, generated by AI itself, and new sources of data.

The library and the student

A student reads thousands of books in a library and learns to write well, without copying any single book. AI companies argue their models learn similarly. Authors reply that the student bought or borrowed books legitimately, while AI companies copied works at massive scale for commercial products without paying. The analogy is at the heart of the legal debate.

Thinking data is free because it is online

Content available online is often protected by copyright. Whether AI companies can use it for training without permission is legally contested, and licensing deals show that data increasingly has a price.

Key takeaways
  • AI models learn from vast amounts of data, making data a key economic input.
  • Authors, artists and publishers have sued AI companies over training on copyrighted work.
  • Anthropic agreed in 2025 to pay 1.5 billion dollars to settle an authors' lawsuit.
  • Licensing deals are creating a market for data, and high-quality data may become scarce.
4 min read

No recording for this one yet - EconReader can read it aloud for you.

Welcome to EconReads

This site is made for visually impaired learners, so our read-aloud reader is already switched on to help you explore hands-free.

You're in control - turn it off any time using the Reader button at the top of the page.

EconReader Ready