Privacy Foundations for Multi-Institutional Scientific Artificial Intelligence
A perspective frames scientific-AI privacy as institutional assurance, flagging scheduler telemetry and instrument side channels.
This perspective treats privacy for multi-institution scientific AI as an assurance problem defined by protected asset, observer, channel, permitted disclosure, guarantee, and evidence. Differential privacy, federated learning, secure computation, trusted execution, and provenance do not cleanly compose across mixed-trust laboratories, access tiers, and autonomous agents. It flags scheduler and telemetry metadata that reveal resource posture, and instrument control loops that leak research strategy through timing and contention on shared accelerators, and lists six research priorities.
- Privacy claims should specify asset, observer, channel, disclosure, guarantee, and evidence.
- DP, federated learning, MPC, and TEEs rarely compose across mixed-trust institutions.
- Scheduler and allocation telemetry can expose an institution's compute posture.
- Shared-accelerator timing and contention can leak research strategy.
Full article199 words · extracted from arxiv.org · click to collapse
Scientific artificial intelligence (AI), spanning foundation models (FMs) to federated data-analysis pipelines, is becoming shared infrastructure across national laboratories, universities, hospitals, and industrial partners. This collaboration creates privacy risks whose natural unit is often an institution's participation, research strategy, or technical capability rather than a single record. Differential privacy (DP), federated learning (FL), secure computation, trusted execution, and provenance each protect parts of the stack, but their guarantees rarely compose across mixed-trust institutions, access tiers, and autonomous agents. This perspective recasts privacy for scientific AI as an assurance problem defined by six elements: protected asset, observer, channel, permitted disclosure, guarantee, and evidence. We demonstrate the framing through a claim register for a composite cross-institutional scenario and use it to assess the model lifecycle. Two of the resulting gaps are specific to leadership-class facilities: scheduler, allocation, and telemetry metadata expose an institution's resource posture, and instrument-attached control loops leak research strategy through timing and contention on shared accelerators. We identify six research priorities: institution-level guarantees, agent-communication privacy, cross-tier information flow, privacy-compatible reproducibility, leadership-scale accounting, and instrument side channels. The contribution is a common form for stating, comparing, and auditing claims whose guarantees otherwise remain fragmented across the scientific AI stack.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.39787