Skip to preprint detailsSkip to PREreviews

PREreviews of HE-Guardrail: A Homomorphic Guardrail Against Jailbreak Attacks for Encrypted Large Language Model Inference

1 PREreview

  1. PREreview by Karan Shah

    Does the introduction explain the objective of the research presented in the preprint?
    Yes
    The introduction makes the problem clear early on: because the server only ever sees ciphertexts, it can't tell when someone sends a jailbreak prompt or when the model produces something harmful, and…
    Read the PREreview by Karan Shah