← Front page

The a16z Show · Saturday, September 5, 2026

Levie Questions Distinguishing Between Data Training Sources

Aaron Levie raised questions about the ethical distinctions between training AI models on public internet data versus outputs from other AI models. He argued that it's difficult to draw a clear ethical line between these two approaches, especially since many individuals did not explicitly consent to their data being used for initial training runs.

The tape

3 quotes
I'm almost deeply on the side of I think it's very hard to make the argument that AI models should be trained on broadly the public internet, but another AI model can't be trained on the outputs of an AI model.
Aaron Levie
Maybe it's in some kind of terms of service from the underlying provider somewhere in the fine print.
Aaron Levie
But like this is a thing that sort of like we are assuming that the general knowledge of the world is going to be trained into these models.
Aaron Levie
Heard on The a16z Show — “Aaron Levie on Why Open AI Wins, published Saturday, September 5, 2026. Heardvine summarizes and quotes with attribution and timestamps, and links to the original everywhere.
Transcribed via Gemini audio transcription · $0.04