Encrypted neural-network evaluation can require approximating an activation function with a polynomial. That approximation works within a chosen input range; values outside the range can grow far faster than the original activation. The Homomorphic Advantage Operator paper examines how reinforcement learning's repeated value updates can push a network into that unstable region.
Abid Mohamed Nadhir and colleagues propose centering the action-value estimates used in temporal-difference targets. Their system combines that operator with reward scaling and regularization. The reported experiments concern numerical stability under specific polynomial and encryption constraints, rather than a general privacy certification for reinforcement-learning systems.
The recursive target creates a stability problem
In supervised learning, a network often learns against fixed labels. A value-based reinforcement learner constructs its target using estimated values for a successor state. An approximation error can therefore affect later targets as training continues.
The authors call the resulting amplification Bellman drift. Their account connects a uniform state-value baseline with pressure on network pre-activations, which must stay inside the polynomial's valid approximation domain. Weight decay and gradient clipping constrain updates, but the paper argues they do not remove that shared baseline from the recursive target.
The practical implementation uses leveled homomorphic encryption. It allocates a finite computation depth to avoid ciphertext bootstrapping. That detail narrows the title's broad fully homomorphic encryption framing: the implemented circuit does not perform arbitrary-depth encrypted computation.
Centering action values before the backup
HAO subtracts the mean action value for a state from each action value. Because the same scalar is subtracted from every component, centering preserves the greedy ranking of that fixed action-value vector. The authors implement this operation with a precomputed linear projection supported by CKKS arithmetic.
Their operator adds no nonlinear multiplicative depth, although the paper notes a plaintext-ciphertext product can consume a rescaling level. Zero nonlinear depth therefore should not be interpreted as zero computational or cryptographic cost.
The full approach also scales rewards on the client side and uses decoupled weight decay with gradient clipping. The paper states that centering and regularization work together in its stabilization design. An operator-only account would omit conditions the authors consider necessary.
Preserving the ordering of one vector also has a limited meaning. Changing the bootstrapping target changes the learning recursion; the fixed-vector ranking property alone does not prove that a trained agent recovers the original optimal policy in every environment. The paper distinguishes ideal row centering from a sample-based successor-state target in its formal analysis.
Experiments span different levels of implementation
The study describes a tabular Markov decision process, CartPole with real TenSEAL CKKS operations, and a logistics-routing benchmark with continuous features. Its abstract reports no polynomial-boundary breaches for HAO across the random seeds evaluated, alongside breaches in comparison configurations.
Those observations are finite experimental results. They do not guarantee that arbitrary networks, reward scales or encryption parameters will remain stable. The relevance of a reproduction depends on retaining the activation approximation, admissible domain and stabilization settings rather than comparing the operator name alone.
The authors also report stability when Gaussian noise of the kind used in DP-SGD is added to clipped gradients. That is a numerical compatibility observation. A claim of differential privacy would additionally require its specified privacy mechanism and accounting; the reported noise experiment alone supplies no general privacy budget.
For an engineering assessment, this paper identifies a concrete interaction between value bootstrapping and encrypted polynomial evaluation. The key questions are whether the chosen network remains within its approximation bounds and whether changing the target preserves the decision quality needed for the particular task.
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!