AIGP · Domain 4
AI Ethics and Responsible AI
About 25% of the exam
Principles landscape
OECD AI Principles
- Inclusive growth, sustainable development, well-being
- Human-centered values and fairness
- Transparency and explainability
- Robustness, security and safety
- Accountability for outcomes
- Plus five recommendations for policymakers
EU HLEG seven requirements
- Human agency and oversight
- Technical robustness and safety
- Privacy and data governance
- Transparency, traceability, communication
- Diversity, non-discrimination and fairness
- Societal and environmental well-being
- Accountability, including auditability
Also on the exam
- ALTAI: self-assessment checklist for the seven
- UNESCO: four values, non-binding
- IEEE 7000: values into requirements
- IEEE 7001: transparency by stakeholder level
- HUDERIA: rights, democracy, rule of law
- Principles-to-practice gap is the critique
Ethics vocabulary
- Beneficence
- do good, deliver benefit
- Non-maleficence
- do no harm, prevent it
- Autonomy
- respect informed, uncoerced choice
- Justice
- fair distribution, no worsened inequity
- Proportionality
- intrusion justified by the objective
- Necessity
- no less intrusive way exists
- Precautionary principle
- uncertainty does not excuse delay
- Dual use
- benefit and misuse from one capability
Governing ethics
- Ethics board: diverse, external voices, clear mandate
- Authority to pause or block deployments
- Escalation route to executives and board
- Participatory design with affected communities
- Accountability shown by RACI, not statements
- Review triggered above a risk threshold
- A simple rule beats needless AI
- Staged release for dual-use capability
Where bias comes from
Data
- Historical bias baked into outcomes
- Representation: groups missing or sparse
- Measurement: proxy target such as cost
- Label bias from inconsistent annotators
- Proxy features track protected attributes
- Feedback loops recycle past decisions
Model
- Learning bias: objective sacrifices small groups
- Aggregation: one model, different populations
- Evaluation bias: benchmark unlike deployment
- Compression degrades rare cases first
- Threshold choice sets group error rates
Deployment
- Used outside the validated context
- Automation bias: humans defer to output
- Problem framing: punitive versus assistive
- Distribution shift after launch
- No monitoring of real outcomes
Arrest data reflects where police patrolled; a model trained on it sends them back there
Fairness metrics
- Demographic parity
- equal positive rate per group
- Equal opportunity
- equal true positive rates
- Equalized odds
- equal true and false positive rates
- Predictive parity
- equal precision across groups
- Calibration
- same score, same real probability
- Individual fairness
- similar people, similar outcomes
- Counterfactual fairness
- outcome unchanged if group flipped
- Four-fifths rule
- selection ratio below 0.8 flags
Fairness trade-offs
- Metrics conflict when base rates differ
- Calibration versus equal error rates: COMPAS
- Choose and justify a metric for the harm
- Fairness-accuracy trade-off decided deliberately
- Parity can approve the unqualified
- Group-specific thresholds may be disparate treatment
- Document the threshold metrics were measured at
- Inferred demographics carry correlated error
- Intersectional slices: small, unstable, spurious
Bias mitigation across the lifecycle
- Framing
- Data
- Pre-processing
- In-processing
- Post-processing
- Monitoring
- Framing
- target variable decides who bears error
- Pre-processing
- reweight, resample, relabel, model agnostic
- In-processing
- fairness constraints inside the objective
- Post-processing
- adjust thresholds on a fixed model
- Measurement
- protected attributes only to test disparity
- Monitoring
- deployer tracks disaggregated outcomes
You cannot fix a disparity you refuse to measure; collect or infer group data solely for that purpose
Law hooks for discrimination
- Disparate treatment
- intentional rule against a group
- Disparate impact
- neutral rule, unequal outcome
- Four-fifths rule
- adverse impact evidence threshold
- Proxy discrimination
- correlated feature stands in
- ECOA, Regulation B
- specific adverse action reasons
- NYC Local Law 144
- annual independent bias audit
- AI Act Article 10
- special data only to detect bias
Transparency and explainability
- Transparency: disclose use, data, purpose, limits
- Explainability: why this particular output
- Interpretability: model understood by design
- Global explains the model, local one decision
- LIME: local surrogate around one instance
- SHAP: contribution value per feature
- Model cards report disaggregated performance
- Datasheets document composition and limits
- Explanations must support contesting the decision
Human oversight and accountability
- Human in the loop
- approves each decision
- Human on the loop
- monitors, can intervene
- Human in command
- decides whether and how to use
- Meaningful review
- time, information, authority to change
- Automation bias
- rubber-stamping the machine
- Contestability
- appeal path with real correction
- Accountability by design
- owners assigned before deployment
- Traceability
- data, version, rationale per output
Harms and manipulation
- Manipulation exploits vulnerability to distort choice
- Dark patterns and consent fatigue
- Engineered emotional dependency for retention
- Engagement optimization amplifies outrage
- Algorithmic management without recourse
- Deepfake intimate imagery harms a real victim
- Quality-of-service harm versus allocative harm
- Children and vulnerable groups need heightened care
- Surveillance: proportionality, necessity, oversight
Responsible AI in operation
- Red teaming: adversarial, recurring, documented
- Staged deployment with gates
- Disaggregated evaluation before release
- Post-market monitoring with re-assessment triggers
- Incident and near-miss reporting
- Appeal path for affected people
- Diverse teams surface blind spots
- Track environmental cost of training
- Publish responsible AI reports
Rapid recall: name the fairness idea
- Same approval rate per group
- demographic parity
- Same TPR and FPR
- equalized odds
- Same TPR only
- equal opportunity
- Scores mean the same everywhere
- calibration across groups
- Similar people, similar outcomes
- individual fairness
- Ratio under 0.8
- four-fifths flag
- Zip code stands in for race
- proxy discrimination
- Explicit rule against a group
- disparate treatment
Reference strip: principles, bias, metrics, oversight, harms
Principles
- OECD five, HLEG seven
- ALTAI checklist, UNESCO values
- Beneficence, non-maleficence, autonomy, justice
- Proportionality and necessity
Bias sources
- Historical, representation, measurement
- Learning, aggregation, evaluation
- Deployment context, automation bias
- Proxies and feedback loops
Metrics
- Parity, equal opportunity, equalized odds
- Calibration conflicts with error balance
- Individual and counterfactual fairness
- Four-fifths rule, 0.8
Oversight
- In the loop, on the loop, in command
- Meaningful review, not rubber stamp
- Contestability and traceability
- Ethics board can block
Harms
- Manipulation and dark patterns
- Dependency, outrage, surveillance
- Deepfakes, dual use
- Vulnerable groups first
Quick exam traps
- Trap: Removing the protected attribute from the features removes the bias
- Trap: A model can satisfy every fairness metric at once
- Trap: Demographic parity is always the right fairness target
- Trap: Calibration across groups proves a tool is fair
- Trap: A reviewer approving 99% of outputs in seconds is meaningful oversight
- Trap: Explainability and interpretability are the same property
- Trap: An ethics board is advisory only and should never block a launch
- Trap: High engagement proves the content serves users' interests
cybercertprep.com · original revision sheet written from the public body of knowledge