The a16z Show · Saturday, September 5, 2026
Aaron Levie raised questions about the ethical distinctions between training AI models on public internet data versus outputs from other AI models. He argued that it's difficult to draw a clear ethical line between these two approaches, especially since many individuals did not explicitly consent to their data being used for initial training runs.
“I'm almost deeply on the side of I think it's very hard to make the argument that AI models should be trained on broadly the public internet, but another AI model can't be trained on the outputs of an AI model.”
“Maybe it's in some kind of terms of service from the underlying provider somewhere in the fine print.”
“But like this is a thing that sort of like we are assuming that the general knowledge of the world is going to be trained into these models.”