Assume the dataset will grow and eventually mimic reality.
How would that happen, exactly?
Stereotypes themselves and historical bias can bias data. And AI trained on biased data will just learn those biases.
For example, in surveys, white people and black people self-report similar levels of drug use. However, for a number of reasons, poor black drug users are caught at a much higher rate than rich white drug users. If you train a model on arrest data, it'll learn that rich white people don't use drugs much but poor black people do tons of drugs. But that simply isn't true.