Developers of large language models such as OpenAI and Anthropic are increasingly interested in internal company data because they face limitations in gathering publicly available information on the web to train neural networks. Demand for data such as instant messaging tools, emails, video conference recordings, and code change history has increased dramatically in recent months, sources said.

Image source: Gemini
Bobby Samuels, chief executive of data broker Protege, said the company’s trading volume has grown from $30 million last year to at least $100 million this year, driven by growing demand. The growing interest in data is also related to the development of artificial intelligence agents that can perform practical daily tasks for users.
Interest in enterprise data stems from the fact that records of real-life interactions between people contain information about how developers solved problems, discussed finances, demonstrated how software worked and performed other tasks. All of these are very useful when training and improving artificial intelligence models. Buyers are primarily interested in data from startups facing bankruptcy or acquisitions.
For example, Warmly, the AI agency startup recently acquired by HubSpot, received four offers worth up to $300,000 for employee work interaction data, including meeting notes and emails. All offers were made after the acquisition agreement was signed, but all were rejected.
As enterprise data transactions grow, the problem of impersonality becomes more acute. Companies that sell artificial intelligence training materials are being forced to pay more attention to removing personally identifiable information. “We follow strict anonymization processes to protect the value of the data and comply with privacy regulationssaid Shub Sinha, CEO of data processing company Integral.
If you find an error, select it with your mouse and press CTRL+ENTER.
