I assumed you want to make person unidentifiable, in which case tone and pitch changes will not erase enough of biometrics.
If you want to make it sound like someone specific, then it's exactly what you said in the last line, you need to train your own spoofing models. Also you sound like you want to do something I can't endorse, so it's good that you don't want to train a model