Structured PREreview of HE-Guardrail: A Homomorphic Guardrail Against Jailbreak Attacks for Encrypted Large Language Model Inference
- Published
- DOI
- 10.5281/zenodo.23289740
- License
- CC0 1.0
- Does the introduction explain the objective of the research presented in the preprint?
- Yes
- The introduction makes the problem clear early on: because the server only ever sees ciphertexts, it can't tell when someone sends a jailbreak prompt or when the model produces something harmful, and most HE work simply assumes the client is honest. The authors then say plainly what they want to do, which is run existing guardrails directly on the encrypted data and use the result to decide, still under encryption, whether the client gets the real answer or a refusal. Figure 1 and the contribution list help a lot here. My only problem is that they call intra-processing guardrails "promising" before showing any results, and their own numbers later don't really back that up.
- Are the methods well-suited for this research?
- Somewhat appropriate
- The general approach makes sense. They pick one pre-processing guardrail and two intra-processing ones, run them under real CKKS encryption, and evaluate them with an established framework (the SEU setup from the SoK paper), standard attack sets, and a standard harm judge. That gives a fair answer to the main question of whether guardrail decisions survive encryption. There are a few gaps, though. Table 1 has no plaintext guardrail column, so you can't tell whether the weak JBShield and GradSafe results come from encryption or from how they were reimplemented. All the attacks are non-adaptive, even though the setting arguably makes adaptive attacks easier. The noise flooding defense is only argued with a toy example, with no analysis of what happens when a client repeats queries. And the latency claim rests on a second GPU without any actual timing numbers.
- Are the conclusions supported by the data?
- Somewhat unsupported
- Some conclusions hold up. The 99.7% agreement figure does back the claim that encryption rarely changes a guardrail's decision, and the results section is fairly honest in places, for example when it notes that the low false-positive rates for JBShield and GradSafe just reflect how permissive they are. The conclusion section goes further than the data allows. It says the encrypted guardrails perform close to their plaintext versions, but the paper never reports plaintext numbers to compare against, only agreement with the authors' own reimplementation. It also says the suppressed response "cannot be recovered" by the client, which is never tested or proven, and repeated queries look like an obvious way around it. The broader message that HE-Guardrail defends against jailbreaks is also hard to square with Table 1. Two of the three guardrails leave the attack success rate basically unchanged, and only Llama Guard makes a real difference.
- Are the data presentations, including visualizations, well-suited to represent the data?
- Somewhat appropriate and clear
- The tables are simple and easy to read. Table 1 puts ASR and pass-guardrail rate side by side for each attack family, which makes it easy to spot that JBShield and GradSafe don't actually move the attack success rate. Table 3 helpfully gives raw rejection counts next to the percentages. Figure 1 is a good overview of the pipeline, and it labels everything with text, so it doesn't rely only on the red/blue color coding. The main problems are things left out rather than things shown badly. Table 1 really needs a plaintext column for each guardrail, since that comparison is central to the paper's claims. Table 2 gives only rough memory estimates. There's no table or figure with actual latency numbers, even though the paper makes a fairly strong claim about zero added latency. A small breakdown of the 0.3% of cases where the encrypted and plaintext decisions differ would also have been useful.
- How clearly do the authors discuss, explain, and interpret their findings and potential next steps for the research?
- Neither clearly nor unclearly
- The results sections mostly restate the numbers in the tables without explaining them. The most important finding, that JBShield and GradSafe block a lot of prompts but barely lower the attack success rate, gets no real interpretation. The authors don't ask why their versions perform so much worse than the original papers report, or what this means for the "intra-processing is promising" argument from the introduction. The future work section is short and covers only multi-turn attacks. That is a fair point, and they explain it reasonably well. But there's no limitations section, and obvious next steps are missing: adaptive attacks, a formal analysis of the noise flooding, checking that client inputs are valid, and post-processing guardrails. The conclusion also reads more positively than the results justify, which makes the overall interpretation feel a bit thin.
- Is the preprint likely to advance academic knowledge?
- Somewhat likely
- The paper's biggest contribution is the problem it raises. Most HE inference work assumes the client is honest, and this paper makes a convincing case that a malicious client with encrypted inputs is a real blind spot. That framing alone is likely to shape follow-up work. There are also some solid technical pieces. The most notable is running GradSafe's backward pass under encryption, which hasn't really been done for a model this size. The observation that the server's white-box access makes intra-processing guardrails a natural fit for HE is also useful, and so is the multi-model serving trick. What keeps it from a higher rating is that the security results are weak for two of the three guardrails and the leakage protection isn't fully worked out. It works better as a starting point for this line of research than as a finished solution.
- Would it benefit from language editing?
- No
- Would you recommend this preprint to others?
- Yes, but it needs to be improved
- Is it ready for attention from an editor, publisher or broader audience?
- Yes, after minor changes
Competing interests
The author declares that they have no competing interests.
Use of Artificial Intelligence (AI)
The author declares that they used generative AI to come up with new ideas for their review.