ZeroHour

Search: “llama”

11 stories

An Open-Source End-to-End FHE Implementation for Privacy-Preserving Llama 3 8B Inference

Odin runs Llama-3-8B fully homomorphic encrypted inference on a single H100 in 366 seconds, a 4.51x speedup over THOR.

Odin is an open-source end-to-end GPU CKKS implementation for privacy-preserving Llama-3-8B inference that co-designs ciphertext packing with model execution. A feature-major cross-layer layout unifies residual connections and layer interfaces, while transient intra-operator layouts serve linear projections and attention, avoiding intermediate repacking of QK^T softmax outputs. Minimax polynomial approximation with input-range control reduces polynomial degree and multiplicative depth for nonlinear ops. With 128-token input, Odin evaluates all 32 Transformer layers on one NVIDIA H100 80 GB in 366.4 s using 58.9 GiB peak memory, versus 1651.9 s for the THOR baseline, a 4.51x speedup.

arXiv cs.CR · 6d agoResearch

SpliTEE: Improving LLM Inference on Trusted Hardware with Differentially Private GPU Outsourcing

SpliTEE splits LLM inference between Intel TDX trusted execution and untrusted GPUs, using differential privacy instead of encryption to protect intermediate representations.

SpliTEE extends split inference to LLMs, running inference partly inside an Intel TDX TEE while masking intermediate inputs sent to untrusted GPUs with differential privacy rather than encryption. The authors show a prompt-reconstruction attack recovers nearly 80% of prompts from unmasked intermediate representations, motivating the masking. A global sensitivity analysis bounds the required DP noise scale, avoiding quantization and keeping models in floating point. The implementation is nearly twice as fast as full CPU-based TDX inference and 5-15 seconds faster than encryption-based Slalom with higher accuracy, evaluated on Llama-3.2-3B and Qwen3-4B.

arXiv cs.CR · 3d agoResearch

Understanding the Privacy-Preserving Potential of HTTP/2 Against Webpage Fingerprinting

Researchers show HTTP/2 features can emulate website fingerprinting defenses like FRONT and Tamaraw with tunable privacy-overhead trade-offs.

An arXiv paper demonstrates that application-layer defenses such as HTTPOS, LLaMA, FRONT, ALPaCA, and Tamaraw can be emulated through HTTP/2 features at both the client and server side, including proactive resource suggestion, multiplexing, and flow control. The authors propose a unified evaluation blueprint that calibrates defense parameters per dataset, combines practical attacks with information-theoretic leakage estimators, and measures overheads to map each defense's privacy-overhead trade-offs.

arXiv cs.CR · 12d agoResearch