AI ethics labels hide misaligned moral reasoning
AI models often agree with human ethical verdicts but for different moral reasons, making label-based alignment tests misleading.
AI models often agree with human ethical verdicts but for different moral reasons, making label-based alignment tests misleading.