--- title: "Distributed Maximum Likelihood Estimation" author: "Balasubramanian Narasimhan" date: '`r Sys.Date()`' bibliography: homomorphing.bib output: html_document: fig_caption: yes theme: cerulean toc: yes toc_depth: 2 vignette: > %\VignetteIndexEntry{Distributed Maximum Likelihood Estimation} %\VignetteEngine{knitr::rmarkdown} \usepackage[utf8]{inputenc} --- ```{r echo=F} knitr::opts_chunk$set( message = FALSE, warning = FALSE, error = FALSE, tidy = FALSE, cache = FALSE ) ``` ## The statistical problem Suppose we have count data $y_1, y_2, \ldots, y_n$ that we model as independent draws from a Poisson distribution with unknown parameter $\lambda$. The maximum likelihood estimate is $\hat{\lambda} = \bar{y}$, obtained by minimizing the negative log-likelihood $$ -\ell(\lambda \mid y) \;=\; -\sum_{i=1}^{n} \log p(y_i \mid \lambda). $$ In R, this is a one-liner using `stats4::mle()`: ```{r} library(stats4) set.seed(17822) y <- rpois(n = 40, lambda = 10) nLL <- function(lambda) -sum(stats::dpois(y, lambda, log = TRUE)) fit0 <- mle(nLL, start = list(lambda = 5), nobs = NROW(y)) summary(fit0) logLik(fit0) ``` ## The privacy constraint Now suppose the same data is **distributed across three sites** — say, three hospitals counting adverse events. None will share its raw counts with the others or with a central aggregator, but they are willing to *jointly* compute the same MLE provided no party learns anything about another party's contribution. To simulate this, partition `y`: ```{r} y1 <- y[1:20] y2 <- y[21:27] y3 <- y[28:40] ``` The negative log-likelihood factorizes additively: $$ -\ell(\lambda \mid y) \;=\; -\ell_1(\lambda \mid y_1) - \ell_2(\lambda \mid y_2) - \ell_3(\lambda \mid y_3) $$ so each site can compute its local term in the clear and only the *sum* of the three local likelihoods needs to travel between parties — and the sum must not reveal the individual addends. ## The protocol We use the **master/worker** topology that distcomp- and DataSHIELD-style federated analyses actually deploy: a star with the master at the center and one independent worker per site. There is no chain and no inter-site communication. In words: 0. The master generates a CKKS context and key pair, distributes the public key to the three workers, keeps the secret key. 1. The master broadcasts the current $\lambda$ to each worker. 2. Each worker computes its local negative log-likelihood $\ell_i(\lambda)$ on its private data, encrypts the result under the master's public key, and returns $E(\ell_i)$ to the master. 3. The master sums the encrypted contributions homomorphically: $E(\ell) = E(\ell_1) \boxplus E(\ell_2) \boxplus E(\ell_3) = E(\ell_1 + \ell_2 + \ell_3)$. 4. The master decrypts $E(\ell)$ to recover $\ell$. 5. The master hands $\ell$ to the optimizer; the protocol repeats for each new guess of $\lambda$ until convergence. This is the realistic shape: each worker independently does its local computation and ships an encrypted summary; the master only sees the encrypted summaries (and, after homomorphic summation, the decrypted total). With a single-decrypter master, the master *could* decrypt individual $E(\ell_i)$ in principle; the cryptographic story strengthens when paired with threshold key generation, where no single party holds the secret key. We will revisit that in the Cox threshold vignette. ## Implementation The computational topology — `Site` (worker), `Master`, the master/worker runner — is the same code as in any other distributed-stats vignette in this package; only the master's *backend* changes. `homomorpheR` exports `make_ckks_master()` that takes an `openfhe.R` `CryptoContext` and `KeyPair`. The master decrypts with the secret key; each site encrypts its own contribution under the public key it is given at setup. The worker class and the `master_aggregate()` runner are backend-agnostic. The per-site negative log-likelihood is the same plain R function it would be in the cleartext case: ```{r} library(homomorpheR) local_nll <- function(data, lambda) { -sum(stats::dpois(data, lambda, log = TRUE)) } ``` ### 1. Generate a CKKS key pair We call `openfhe.R` functions qualified rather than attaching `openfhe.R`. There is one `encrypt()` and one `decrypt()`: they are `openfhe.R`'s generics, and `homomorpheR` registers its methods on them and re-exports them. ```{r} cc <- openfhe.R::fhe_context("CKKS", multiplicative_depth = 1L, scaling_mod_size = 50L, batch_size = 8L) keys <- openfhe.R::key_gen(cc) ``` ### 2. Build workers and master ```{r} worker1 <- make_worker("S1", data = y1, contribution_fn = local_nll) worker2 <- make_worker("S2", data = y2, contribution_fn = local_nll) worker3 <- make_worker("S3", data = y3, contribution_fn = local_nll) master <- make_ckks_master("Master", crypto_context = cc, keypair = keys) set_workers(master, list(worker1, worker2, worker3)) ``` ### 3. Run `mle()` through the encrypted channel ```{r} fit1 <- mle(function(lambda) master_aggregate(master, lambda), start = list(lambda = 5)) summary(fit1) logLik(fit1) ``` The CKKS-based estimate differs from the cleartext estimate `fit0` by `r formatC(abs(coef(fit1)[["lambda"]] - coef(fit0)[["lambda"]]), format = "e", digits = 1)`. No site ever revealed its individual counts to any other party. ## Beyond MLE For Poisson MLE the protocol uses only additions, so the additive Paillier scheme would also suffice. CKKS saves the integer encoding of real numbers that Paillier needs. The same holds for stratified Cox regression (`vignette("cox")`), where the sites' partial log-likelihoods are summed. Multiplying encrypted values is needed only when the computation itself runs on encrypted data, as in the sigmoid of `vignette("encrypted-regression")`. ## CAVEAT This is a teaching example. In production you would want a real communication transport, threshold key generation so no single party holds the full secret, and persistent serialization at site boundaries. The point of this vignette is not the deployment story but the *structure* of a privacy-preserving distributed computation built on homomorphic encryption.