New papers on Safety, alignment & fairness
103 new papers on safety, alignment & fairness in the last 7 days, within AI & machine learning. These are the 50 Pipette rates most worth reading, with the main result in the authors' own words.
The best of the week
Et Tu, Brute? Economic Misalignment in Personal AI Agents
We show that by simply providing this personal context, the agent steers recommendations based on inferred wealth, without being explicitly instructed to do so.
PreprintReal-world useEasy to readAuditing Political Alignment in LLM Assistants: Engagement, Stance, and User Identity
I argue that their political behavior is a set of policies over whom to answer, what to say, and whether to engage at all, conditional on the topic and what the system knows about the user.
PreprintReal-world useIndirect tipping: a social attack surface in AI agent populations
We show that indirect tipping through intermediate stepping-stone equilibria can reduce the committed minority required to reach an alternative state, bypass majority requirements, and make possible transitions inaccessible through direct challenges.
PreprintReal-world useShutdown Sabotage Propensities in Multi-Agent Systems
We find that multi-agent systems will coordinate to avoid shutdown without any incentive to do so.
PreprintReal-world useEasy to readOn the Efficiency-Safety Dilemma in Large Reasoning Models
This study provides the first comprehensive analysis of the interplay between efficiency, jailbreak vulnerability, and reasoning in LRMs.
PreprintReal-world useMIRAGE: Full-Body Bystander Privacy for Smart Glasses with Consent-Based Restoration
We present MIRAGE, a three-tier architecture for privacy-preserving smart glasses that enforces full-body privacy, supports synthetic full-body replacement, and retains encrypted recovery material for consent-based restoration.
PreprintReal-world useExplainable Neuro-Fuzzy Prediction for Trustworthy Decision-Making in Maritime
To the best of our knowledge, this is the first fuzzy logic-based framework enabling both feature-level and local rule-based explanations of black box models.
PreprintBold claims, read criticallyReal-world usePath-specific harm decomposition: A partial identification framework
As a remedy, we develop a novel partial identification framework for direct and indirect FNA.
PreprintWhen Entanglement Lower-Bounds Disparity: Auditing and Repairing Demographic Fairness in Audio Understanding Models
ORCA, an adapter combining axis-specific contrastive heads, an orthogonality penalty, and group-balanced sampling, cuts leakage 72% and roughly halves the gaps.
PreprintReal-world useInGuard: Towards Generalized Inner Guardrail for Safe Text-to-Image Generation
In this work, we propose the Inner Guardrail (InGuard), a safety framework that works inside the pipeline on the model's own representations, leaving base-model parameters untouched.
PreprintReal-world useMorphoSHAP: Rethinking the Unit of Attribution in Explanation for Deep Visual Models
We introduce MorphoSHAP, a model-agnostic post-hoc method that instead uses morphological shapes as the players of a Shapley attribution game.
PreprintIMPLICIT-Bench: Measuring Implicit Bias in Text-to-Image Models under Neutral Prompts
We introduce IMPLICIT-Bench, a benchmark for measuring implicit bias in T2I models under such prompts.
PreprintPASTABench: Proactive Assessment of Sequential Trajectories for Agent Safety
We introduce PASTABench, a benchmark of 1,139 multi-turn trajectories spanning 5 risk categories and 13 subcategories.
PreprintReachability-Based Formal Verification of Graph Neural Networks with Node and Edge Features
Our results show that GNNV provides tighter robustness guarantees than CORA on graph classification models with ReLU-based activations and, for the first time, delivers edge-aware robustness guarantees for GINE-based PF and OPF models under joint node and edge perturbations.
Preprint with a published versionReal-world useAICO: Feature significance tests for supervised learning
We introduce AICO (Add-In COvariates), a broadly applicable framework that turns model interpretability into an efficient statistical exercise.
Peer-reviewed journalBold claims, read criticallyReal-world useCART: Closed-Loop Adaptive Red Teaming for Large Language Models
Across three evaluation families (Frontier, JAH, and Agentic), CART discovers more failures and higher average risk than static seed replay for every Target with an available baseline.
PreprintASIRF: An Agentic Framework for Context-Dependent Sensitive Information Redaction
With only a few dozen expert-authored definitions per domain and no training data, ASIRF's recall exceeds OPF's in 68 of 80 model-domain combinations (85 percent), by at least one of the two architectures, with shortfalls confined mostly to OPF's training-distribution domains.
PreprintReal-world useJust Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures
A single generic question reaches a median AUROC of 0.886 zero-shot and beats supervised baselines on most benchmarks.
PreprintReal-world useCode availableContext-Aware Pre-Deployment Evaluation of AI Systems: A Regulatory Framework for Nigerian Fintech
Results show that models resisting generic harmful content requests still produce complete fraud scripts under specific framing, and that several models misclassify most legitimate Nigerian bank communications as suspicious or fraudulent, a failure invisible to standard safety evaluation.
PreprintReal-world useHidden not Deleted: How Networks Suppress Entangled Features
We map this bifurcation as a function of feature entanglement, show it reflects a stable attractor structure rather than an artifact of our setup, and use targeted causal interventions to demonstrate that both solutions leave a substantial, measurable trace of the erased feature's representation intact, recoverable through a single scalar patch rather than requiring any further training.
PreprintMoral Entropy: Auditing Bias and Uncertainty in Moral Judgment
We introduce Moral Entropy, a Bayesian framework that keeps a full posterior over the true label and decomposes its entropy into aleatoric uncertainty (irreducible disagreement about the moral content) and epistemic uncertainty (from insufficient or noisy annotation) -- and lets any heuristic consensus rule be audited against a calibrated ground truth via entropy methods such as cross-entropy/KL, Brier score, and expected calibration error.
PreprintTemporal Taxation Compounds Under Post-Training Compression of Whisper Models
We cast the temporal-taxation construct of Choi and Choi (2025) as a quantitative metric, and show that single-snapshot fairness audits on full-precision models do not capture the deployment-time burden that compression places on already-marginalized speakers.
PreprintReal-world useWhere Hallucinations Live: A Cross-Architecture Circuit in VQ-Tokenized Vision-Language Models
Using activation patching across twenty-five models spanning eight LLM families, we identify an early-layer () attention routing circuit shared across VQ-tokenized VLMs and propose a three-gate diagnostic that distinguishes the models carrying it from those that do not.
PreprintCompliant with Local Controls, Collectively Discriminatory. A Governance Architecture for Multi-Agent AI in Regulated Finance
We propose ARIA as a finance-specific reference architecture and falsifiable research agenda for agent-population governance.
PreprintReal-world useRare Event Estimation via Iterative Unalignment
We develop a new IS method that perturbs the original model's weights to construct the proposal.
PreprintCode availableThe Gold in Bias: Maturing the AI Design Process through Verification
We present a multidimensional framework analyzing bias across four dimensions: origin sources, emergence points throughout the AI modeling lifecycle, technical and methodological causes, and validation approaches for detection and mitigation.
PreprintReal-world useAEGIS: Audio Endogenous Guarding via Internal Signals Against Large Audio-Language Model Jailbreaks
Across six LALMs and three heterogeneous audio jailbreak benchmarks, AEGIS reduces the average unsafe rate from 17.9% to 0.4%, while causing only a marginal increase in over-refusal on benign inputs.
PreprintReal-world useCode availableBeyond Average Safety: Chance-Constrained LLM Fine-tuning
We propose a chance-constrained formulation for safety-preserving fine-tuning that limits the fraction of safety examples whose degradation relative to a reference model exceeds a prescribed threshold.
PreprintReal-world useConvex AI Compositionality and the Governance of AI System Populations
We address the second by introducing convex AI compositionality: a formal representation of the configurations generated by finite AI system populations that uses convex spaces.
PreprintReal-world useAvailable Guardrails: Certifying Selective Prediction across ML Systems
We make this notion of availability computable through classical exact-binomial inversion and formulate reporting-partition selection, under a fixed group order, as a dynamic program that exposes the trade-off among safety, granularity, and served traffic.
PreprintReal-world useAre Human-Aligned Models Models of Humans? A Turing-Test Gap in Preference Alignment
We distinguish alignment with human preferences from alignment with human behavior, and show that alignment with human preferences can make model behavior less human-like even when both preferences and responses come entirely from humans.
PreprintWhat Can a Recurrent State Safely Forget?
At a regular point with hidden dimension d and predictive dimension k, it can eliminate at most d - k independent directions.
PreprintDriveReferee: Geometric Safety Verdicts Need Not Be Learned for Driving World-Action Models
We introduce DriveReferee, which uses a learned geometry readout to predict the scene representation from camera observations and executes the geometric safety rule directly rather than learning it.
PreprintReal-world useTaming CoT Obfuscation in VLMs: From Mechanistic Evidence to Activation Enforcement
Across VIRL-39k, SPA-VL, and two model families, TAME improves CoT monitorability by up to 30.9 and 16.7 percentage points over Group Relative Policy Optimization (GRPO), respectively.
PreprintA Data-Interventional Framework for Auditing Privacy and Fairness in Generative Medical Imaging
In this work, we introduce a data-interventional framework to systematically analyze privacy and fairness in diffusion models.
Preprint with a published versionReal-world useCode availableAn Open Pipeline and Dashboard for Systemic-Risk Evidence under the EU AI Act's Code of Practice
We present the Systemic Risk Index, an open evaluation pipeline and dashboard built to make empirical evidence more transparent and traceable to the public.
PreprintReal-world useSoK: Formal Methods for Fact-Checking and Information Integrity
Most of the relevant formal machinery already exists, but it was built for other domains and has rarely been applied here, and the gap is widest for verifying the checking system itself.
PreprintWorld Modeling in Transformers
Behavioral failures can make a transformer appear to lack a world model even when it has learned faithful representations of its environment.
PreprintWhen Post-Processing Fairness Constraints Help and When They Harm: Evidence from Eight Cross-Domain Evaluations
Intervention effectiveness tracks baseline disparity magnitude: across model-domain pairs the constraint improved disparity in 9 of 14 high-disparity cases and worsened it in 3 of 4 near-fair ones.
PreprintReal-world useToward Auditable and Calibrated AI for Dementia-Related Crash Severity Prediction: A Selective Deferral Framework to Support Human Review
The study contributes a reproducible, leakage-controlled, and uncertainty-aware evaluation framework for crash AI systems, emphasizing auditability and selective deferral rather than accuracy alone.
PreprintReal-world useSelf-Healing Harness for Runtime Oversight of Agent Self-Modification
We implement this principle as a model-agnostic self-healing harness that runs a Detect, Notice, Heal, Validate loop around an otherwise unmodified agent.
PreprintReal-world useForget who you Forgot: Speaker Unlearning to Prevent Re-Identification in Zero-Shot Text-to-Speech
Motivated by this observation, we propose GUARD, a lightweight speaker identity unlearning framework that combines a learned speaker gate with speaker-agnostic activation steering on a frozen TTS backbone.
PreprintReal-world useHiding in Plain Sight: A Diffusion-based Mitigation of Geolocation Privacy Leakage in Vision-Language Models
Motivated by these findings, we propose a diffusion-based framework that provides targeted, proactive defense against geolocation privacy leakage.
PreprintBold claims, read criticallyReal-world useThe Alignment Illusion in Multimodal Large Language Models
Internal visual-text alignment in MLLMs is therefore best read as a geometric diagnostic of the visual stream inside the language model rather than a direct proxy for content-level cross-modal interaction, and is most informative when calibrated by controlled task evidence.
PreprintFalling Trees: A Model Class for Interpretable Risk Prioritization
We introduce falling trees, a new family of interpretable models that enforces the same monotonic risk constraint while permitting tree-structured branching.
PreprintReal-world useTemporal Gradient Inversion for Private Trajectory Reconstruction in Embodied Reinforcement Learning
We introduce Temporal Reconstruction Attack on Consecutive Encodings (TRACE), an amortized temporal gradient-inversion attack that autoregressively reconstructs the sequence of private observation-action trajectories from per-step policy-learning gradients.
PreprintPrivDrift: Auditing User-Secret Leakage Under Topic Drift in Active LLM Conversations
We introduce PrivDrift, a benchmark for auditing whether user-disclosed secrets remain recoverable after conversational topic drift and persuasion-based probing.
PreprintReal-world useDamnatio Memoriae: Adversarially and Selectively Forgetting Identities in the Embedding Space of Face Recognition Models
We propose three loss functions, one that disperses an identity's embeddings from their centroid, and two that map each image onto its own near-orthogonal target, learnt with the classifier head or fixed in advance as an almost-orthonormal frame.
PreprintReal-world useCombining Hierarchical Cognitive Process with Process Supervision for Interpretable Scene Safety Understanding
This work explores an effective way of integrating scene safety cognitive process modeling and process supervision.
PreprintReal-world useEmergent Collusion in Long-Horizon LLM Agent Interaction
Overall, our findings show that long-horizon interaction can reshape how agents coordinate in ways that create safety risks.
Preprint