Data Science for Electron Microscopy
Week 9: Unsupervised learning, autoencoders & latent spaces

Prof. Dr. Philipp Pelz

FAU Erlangen-Nürnberg

Institute of Micro- and Nanostructure Research

FAU Logo IMN Logo CENEM Logo ERC Logo Eclipse Logo

Recap: Week 8 and today’s question

  • Week 8: small data — augmentation, transfer, sim-to-real, active labelling and self-supervised pre-training (SimCLR, MAE, DINO), judged by linear probe vs fine-tune.
  • The remaining gap: 10 000 unlabelled EELS spectra or a 4D-STEM scan and no downstream label at all. We want phases, anomalies and order parameters directly from the data.
  • Today: clustering (k-means, GMM/EM), compression (autoencoders), probabilistic latent spaces (VAE, rVAE) — and how to read latent spaces without fooling ourselves.

Learning outcomes

  1. State k-means (Lloyd, k-means++) and its failure modes on EM data.
  2. Derive the E- and M-steps of EM for a GMM; choose \(K\) with BIC.
  3. Explain linear AE = PCA and why denoising AEs work on low-dose spectra.
  4. Derive the ELBO, explain reparameterisation and the Gaussian KL, diagnose posterior collapse.
  5. Explain how an rVAE separates rotation/translation from content.
  6. Read t-SNE/UMAP and latent maps critically; build denoise → encode → cluster → map → validate.

Road map and self-study

  • Road map: landscape (1) · k-means (2) · GMM & EM (3) · manifold hypothesis & autoencoders (4) · denoising AE, latent space & anomalies (3) · VAE: ELBO, reparameterisation, β (6) · rVAE (1) · how latent spaces mislead (3) · 4D-STEM/EELS latent maps (1) · wrap-up.
  • Self-study: notebooks/week09_autoencoder_vae.ipynb — denoising AE vs PCA, anomaly detection, VAE, GMM + BIC, t-SNE sweep, β-sweep. CPU, ≈1 min.

The unsupervised learning landscape

Four families: clustering, dimensionality reduction, density estimation, generative models.

  • Most EM data arrives unlabelled: clustering proposes phases, compression speeds up maps, anomaly detection flags bad regions — at zero annotation cost.

K-means: objective and Lloyd’s algorithm

Assign \(N\) spectra \(\mathbf{x}_i\in\mathbb{R}^d\) to \(K\) groups \(C_k\) by minimising the within-cluster sum of squares

\[ J = \sum_{k=1}^{K} \sum_{i \in C_k} \|\mathbf{x}_i - \boldsymbol\mu_k\|^2. \]

  1. Assign: each spectrum → nearest centroid.
  2. Update: each centroid → mean of its spectra.
  • Neither step increases \(J\) → converges, but only to a local minimum.

Lloyd’s algorithm on a toy 2-D dataset (stars = centroids).

K-means in practice: initialisation, \(K\), failure modes

  • k-means++: seed centroids with probability \(\propto D(\mathbf{x})^2\); use 5–10 restarts, keep lowest \(J\).
  • Choosing \(K\): no likelihood → heuristics: elbow of \(J(K)\) and mean silhouette score; final check = domain knowledge.
  • Failure modes: assumes spherical, similar-sized clusters; dominated by high-intensity features (zero-loss peak) and outliers; hard labels for mixed boundary spectra.
  • Fix: normalise spectra, or better cluster compressed codes (PCA scores, AE latents) Bruefach, Alexandra et al., (2023), doi:10.1063/5.0130546.

GMM: soft clustering for overlapping phases

\[ p(\mathbf{x}) = \sum_{k=1}^{K} \pi_k \,\mathcal{N}(\mathbf{x};\boldsymbol\mu_k,\Sigma_k), \quad \sum_k \pi_k = 1. \]

  • Each phase = Gaussian cloud around a representative spectrum \(\boldsymbol\mu_k\) with covariance \(\Sigma_k\) → ellipsoidal clusters.
  • Each spectrum gets a responsibility vector \(\gamma_{ik}=P(k\mid\mathbf{x}_i)\) instead of a hard label.

k-means (hard) vs GMM (soft) on the same data; boundary points carry mixed probability.

The latent-variable view and the EM algorithm

Hidden label \(c_i\): \(p(\mathbf{x}_i, c_i{=}k)=\pi_k\,\mathcal{N}(\mathbf{x}_i;\boldsymbol\mu_k,\Sigma_k)\). \(\log\sum_k(\cdot)\) → no closed form. EM alternates:

E-step (Bayes’ rule): \[ \gamma_{ik}=\frac{\pi_k\,\mathcal{N}(\mathbf{x}_i;\boldsymbol\mu_k,\Sigma_k)}{\sum_j \pi_j\,\mathcal{N}(\mathbf{x}_i;\boldsymbol\mu_j,\Sigma_j)} \]

M-step (weighted averages), \(N_k=\sum_i\gamma_{ik}\): \[ \boldsymbol\mu_k=\tfrac{1}{N_k}\textstyle\sum_i\gamma_{ik}\mathbf{x}_i,\quad \Sigma_k=\tfrac{1}{N_k}\sum_i\gamma_{ik}(\mathbf{x}_i-\boldsymbol\mu_k)(\mathbf{x}_i-\boldsymbol\mu_k)^{\top},\quad \pi_k=\tfrac{N_k}{N} \]

EM on Old Faithful: initialisation, iterations, convergence. From Bishop, Christopher M. (2006), Fig. 9.8.

EM guarantees and choosing \(K\) with BIC

  • Monotone: no iteration decreases \(\sum_i\log p(\mathbf{x}_i;\theta)\) — but only a local maximum → k-means init, several restarts Bishop, Christopher M., (2006).
  • k-means = hard-assignment limit of EM (\(\gamma_{ik}\in\{0,1\}\), shared \(\sigma^2 I\), \(\sigma\to0\)).
  • Degeneracy: \(\Sigma_k\to0\) on one spectrum → regularise (reg_covar), diagonal \(\Sigma\), or fit in a latent space.
  • GMM has a likelihood → Bayesian information criterion Schwarz, Gideon, (1978): \[\mathrm{BIC}(K)=-2\log p(\mathcal D;\hat\theta_K)+p_K\log N\]
  • Notebook: GMM on the 2-D VAE latent → BIC minimum at \(K=4\), ARI = 1.00.

The manifold hypothesis: why compression works

Data on a curved low-dimensional surface; PCA fits a flat plane, a nonlinear AE follows the curve.

  • A spectrum is set by a few physical parameters (phase, oxidation state, thickness) — yet we measure 1024 channels. The scree-plot elbow (Week 3) estimates this dimension.

Autoencoder: encoder–bottleneck–decoder

1024-channel spectrum → 8-D latent code \(\mathbf{z}\) → reconstruction.

\[ \mathcal{L}(\theta, \phi) = \frac{1}{N}\sum_{i=1}^{N} \bigl\|\mathbf{x}_i - g_\theta(f_\phi(\mathbf{x}_i))\bigr\|^2 \qquad\text{— no labels; the input is the target.} \]

Linear autoencoder = PCA

  • Linear \(f_\phi(\mathbf{x})=W_e\mathbf{x}\), \(g_\theta(\mathbf{z})=W_d\mathbf{z}\), MSE loss → optimum spans the top-\(k\) PCA subspace Baldi, Pierre et al., (1989).
  • Nonlinearity is what lets an AE follow curved spectral manifolds.
  • Always report the PCA baseline at the same \(k\).

Linear AE (= PCA) vs nonlinear AE on a curved manifold.

Denoising autoencoders for low-dose spectra

Train on corrupted inputs, clean targets Vincent, Pascal et al., (2008): \[\mathcal{L} = \frac{1}{N}\sum_i \|\mathbf{x}_i - g_\theta(f_\phi(\tilde{\mathbf{x}}_i))\|^2\]

  • Noise-specific features cannot help reconstruct the clean signal → only noise-robust features are encoded.
  • Notebook: AE MSE ≈ 5× lower than PCA (\(k\)=4).

Synthetic Fe-L₂₃ spectra at 50 e⁻/channel: clean, noisy, PCA and AE reconstructions.

The latent space: a learned coordinate system

AE latent codes of 400 synthetic spectra (4 phases): islands without labels; stars = k-means centroids.

  • Cluster the \(k\)-D codes, not a 2-D picture of them. Elongated clusters = continuous parameters; isolated points = anomalies.

Anomaly detection: reconstruction error as a score

Train on normal spectra only; unseen phase D reconstructs ~25× worse and is flagged (left); error map (right).

From AE to VAE: the latent space has holes

Left: AE latent with gaps — sampling there decodes to nonsense. Right: VAE latent, pulled toward \(\mathcal{N}(0,I)\), space-filling.

VAE: encoder outputs a distribution, prior on \(\mathbf{z}\)

Decoder (generative): \(\ \mathbf{z}\sim\mathcal{N}(\mathbf{0},I),\quad \mathbf{x}\sim p_\theta(\mathbf{x}\mid\mathbf{z})\)

Encoder (inference): \(\ q_\phi(\mathbf{z}\mid\mathbf{x})=\mathcal{N}\big(\boldsymbol\mu_\phi(\mathbf{x}),\operatorname{diag}\boldsymbol\sigma^2_\phi(\mathbf{x})\big)\)

  • The encoder approximates the intractable posterior \(p_\theta(\mathbf{z}\mid\mathbf{x})\).
  • The union of all \(q_\phi\) tiles the prior → sample \(\mathbf{z}\sim\mathcal{N}(\mathbf{0},I)\) and decode Kingma, Diederik P. et al., (2014).
  • GMM with a continuous latent and a neural likelihood.

Per-input posteriors (ellipses) tile the prior (dashed rings). MFML Unit 11.

The ELBO: derivation in three moves

We want \(\log p_\theta(\mathbf{x})=\log\int p_\theta(\mathbf{x}\mid\mathbf{z})\,p(\mathbf{z})\,d\mathbf{z}\) — intractable for a neural decoder.

  1. Multiply and divide by \(q_\phi(\mathbf{z}\mid\mathbf{x})\); 2. Jensen (\(\log\) is concave); 3. factor \(p(\mathbf{x},\mathbf{z})=p(\mathbf{x}\mid\mathbf{z})p(\mathbf{z})\): \[ \log p_\theta(\mathbf{x})=\log\mathbb{E}_{q_\phi}\!\Big[\tfrac{p_\theta(\mathbf{x},\mathbf{z})}{q_\phi(\mathbf{z}\mid\mathbf{x})}\Big] \;\ge\;\underbrace{\mathbb{E}_{q_\phi(\mathbf{z}|\mathbf{x})}\big[\log p_\theta(\mathbf{x}\mid\mathbf{z})\big]}_{\text{reconstruction}}-\underbrace{\mathrm{KL}\big(q_\phi(\mathbf{z}\mid\mathbf{x})\,\|\,p(\mathbf{z})\big)}_{\text{prior matching}}=\mathrm{ELBO}. \]
  • Gap \(=\mathrm{KL}\big(q_\phi(\mathbf{z}\mid\mathbf{x})\,\|\,p_\theta(\mathbf{z}\mid\mathbf{x})\big)\ge 0\): a better encoder gives a tighter bound.
  • Loss \(=-\mathrm{ELBO}\); reconstruction term = noise-model NLL: Gaussian → \(\tfrac{1}{2\sigma^2}\|\mathbf{x}-\hat{\mathbf{x}}\|^2\) (MSE), Poisson counts → Poisson NLL.

The reparameterisation trick

  • We cannot back-propagate through a sample \(\mathbf{z}\sim q_\phi\).
  • Move the randomness off the parameter path: \[ \mathbf{z}=\boldsymbol\mu_\phi(\mathbf{x})+\boldsymbol\sigma_\phi(\mathbf{x})\odot\boldsymbol\epsilon,\qquad \boldsymbol\epsilon\sim\mathcal{N}(\mathbf{0},I). \]
  • Sampling becomes differentiable → train with ordinary SGD/Adam.

Left: no gradient path through the random node. Right: external noise \(\boldsymbol\epsilon\); red = \(\nabla_\phi f\). From Kingma, Diederik P. et al. (2019).

Closed-form KL and the 10-line training step

For \(q=\mathcal{N}(\boldsymbol\mu,\operatorname{diag}\boldsymbol\sigma^2)\), \(p=\mathcal{N}(\mathbf{0},I)\): \[ \mathrm{KL}(q\,\|\,p)=\tfrac12\sum_{j=1}^{k}\big(\mu_j^2+\sigma_j^2-\log\sigma_j^2-1\big) \]

No Monte Carlo for the KL — only the reconstruction term uses the sampled \(\mathbf{z}\).

mu, logvar = encoder(x_noisy)          # predict log σ² → σ > 0 for free
eps = torch.randn_like(mu)
z = mu + torch.exp(0.5 * logvar) * eps # reparameterisation
x_hat = decoder(z)

recon = ((x_hat - x_clean)**2).sum(1) / (2 * SIGMA2)
kl = 0.5 * (mu**2 + logvar.exp() - logvar - 1).sum(1)
loss = (recon + beta * kl).mean()      # = −ELBO (β = 1)
loss.backward(); opt.step()

β-VAE, posterior collapse and inactive dimensions

  • β-VAE: recon \(+\,\beta\,\)KL Higgins, Irina et al., (2017) — larger β: smoother latent, blurrier output.
  • Posterior collapse: KL → 0, decoder ignores \(\mathbf{z}\).
  • Unused dimensions switch off. Active dims depend on β and decoder variance — not a count of physical degrees of freedom.

Notebook: 4-D VAE; β = 0.1 keeps four dims active, β = 1 and 4 keep one.

Notebook: a denoising VAE on the Fe-L₂₃ spectra

2-D denoising VAE: (a) −ELBO terms, (b) phase latents inside the prior’s 2σ disc, (c) spectra decoded from prior samples.

Recon MSE: noisy 0.00449 · PCA 0.00205 · AE 0.00038 · VAE (\(k\)=2) 0.00017 · ARI 1.00.

rVAE: separating rotation and translation from content

Problem: atom-centred STEM patches differ by rotation and sub-pixel shift; a plain VAE spends its latents on orientation.

  • Latent split: \(\mathbf{z}=(\theta,\Delta\mathbf{r},\mathbf{z}_{\text{content}})\) Bepler, Tristan et al., (2019); Kalinin, Sergei V et al., (2021).
  • Coordinate-based decoder: \(\hat{I}(\mathbf{r})=g_\psi\big(R(\theta)\,\mathbf{r}+\Delta\mathbf{r},\ \mathbf{z}_{\text{content}}\big)\).
  • Rotation/shift act on the pixel grid by construction → \(\mathbf{z}_{\text{content}}\) is orientation-free.
def decode(coords, theta, dr, z):   # coords: (n_pix, 2)
    c = coords @ rot(theta).T + dr  # transform the grid
    h = coord_mlp(c) + latent_mlp(z)
    return out_mlp(h)               # intensity per pixel

Invariance built into the architecture — as in CNNs (Week 7) and GNNs (Week 10).

t-SNE perplexity sweep: same codes, four maps

The same 400 4-D AE latent codes embedded with t-SNE at perplexity 2, 5, 30, 100; ρ = rank correlation of inter-phase distances with the 4-D latent.

How latent spaces mislead — and a checklist

Traps

  • Distance / size: map gaps and areas are not physics.
  • Narrative: humans see blobs even in noise.
  • “Axis means X”: correlation ≠ encoding.
  • Nuisance: latents follow thickness, dose, drift.

Checklist

  • PCA/NMF baseline at the same dimension.
  • Cluster and measure in \(\mathbf{z}\); 2-D maps only for seeing.
  • Sweep perplexity/n_neighbors × seeds; keep what survives.
  • Decode along axes; linear probe on a held-out specimen split.
  • Map back to real space.

Confident is not correct: the interface trap

Notebook: (a) BIC picks K = 4; (b) GMM ellipses and the latent path of Fe₂O₃/FeO mixtures; (c) responsibilities along the path.

A 50/50 Fe₂O₃/FeO interface pixel is labelled “Fe₃O₄” with \(\gamma = 1.00\). A responsibility is a probability under the fitted mixture, not physical truth — check spatial context, reference spectra, linear unmixing (Week 3).

Latent maps for 4D-STEM and EELS

Synthetic 4D-STEM scan (two grains, bending, thickness gradient) → 3 latent coordinates per probe position. \(z_1\) separates grains; \(z_2\) tracks thickness — a nuisance; GMM recovers the boundary (white).

  • Pipeline: denoise → encode → cluster → map → validate against a physical observable (reference fit, virtual detector image, simulation); train/test on different regions.

Notebook summary: Week 9 key results

  • Denoising AE vs PCA (\(k\)=4): MSE 0.00038 vs 0.00205; latent silhouette 0.73 vs 0.63.
  • Anomalies: unseen phase D → 25× higher error; 100/100 flagged.
  • VAE (2-D): MSE 0.00017, KL ≈ 3.4 nats; GMM + BIC → \(K\) = 4, ARI 1.00 — but interface spectra → Fe₃O₄ with \(\gamma\) = 1.00.
  • t-SNE sweep: ρ = 0.26 / 0.54 / 0.83 / 0.77 at perplexity 2 / 5 / 30 / 100.
  • β-sweep: active dims 4 → 1 → 1 for β = 0.1 → 1 → 4.
  • Open In Colab notebooks/week09_autoencoder_vae.ipynb

Summary

  • Clustering: k-means = coordinate descent on within-cluster variance (hard, spherical). GMM + EM = soft, ellipsoidal, with a likelihood → BIC. k-means is the hard limit of EM.
  • Autoencoders: bottleneck → compact code; linear AE = PCA; denoising AEs learn noise-robust features.
  • VAE: loss = −ELBO = noise-model NLL + KL to \(\mathcal{N}(0,I)\); reparameterisation makes it trainable; β switches dimensions off.
  • rVAE: rotation/translation in a coordinate decoder → orientation-free content latents.
  • Latent spaces mislead unless you cluster in \(\mathbf{z}\), sweep, decode, probe honestly, and map back to real space.

Must-know takeaways

  1. k-means minimises \(\sum_k\sum_{i\in C_k}\|\mathbf{x}_i-\boldsymbol\mu_k\|^2\); local minimum → k-means++ and restarts.
  2. GMM: E-step \(\gamma_{ik}\propto\pi_k\mathcal{N}(\mathbf{x}_i;\boldsymbol\mu_k,\Sigma_k)\); M-step = weighted means, covariances, fractions; \(K\) by BIC.
  3. A linear AE with MSE loss spans the PCA subspace; nonlinearity goes beyond PCA.
  4. \(\log p(\mathbf{x})\ge\mathbb{E}_q[\log p(\mathbf{x}\mid\mathbf{z})]-\mathrm{KL}(q(\mathbf{z}\mid\mathbf{x})\|p(\mathbf{z}))\); gap = KL to the true posterior.
  5. Reparameterisation \(\mathbf{z}=\boldsymbol\mu+\boldsymbol\sigma\odot\boldsymbol\epsilon\) makes sampling differentiable.
  6. t-SNE/UMAP: clusters that survive a sweep are evidence; positions, gaps and sizes are not.

Outlook: Week 10 — Attention, transformers & graph networks for EM

  • Today: spectra and patches as flat vectors (MLP) or grids (CNN). The rVAE hinted that where things are and how they relate matters.
  • Attention: queries, keys, values; vision transformers tokenise images and diffraction patterns; MAE = autoencoder with a transformer encoder.
  • Graphs: atom columns as nodes, message passing with permutation invariance — CNN = grid graph, GNN = sparse graph, transformer = full graph.
  • Explainability: attention ≠ explanation, just as a latent axis ≠ a physical variable.

Continue

Backup slides

Material moved out of the 90-minute lecture path; not examined beyond what the main slides cover.

K-means vs GMM: comparison table

K-means GMM
Assignment hard (one cluster) soft (probability vector)
Cluster shape spherical ellipsoidal (full \(\Sigma\))
Speed fast slower (covariance update)
Boundary handling abrupt smooth
Choosing \(K\) elbow + silhouette BIC (penalised log-likelihood)
High-dim risk Euclidean dominance covariance explosion (\(d^2\) params)

Rule of thumb: k-means for fast exploration; GMM when clusters overlap, vary in shape, or you need probabilities for downstream decisions. In both cases, prefer clustering latent codes over raw spectra Murphy, Kevin P., (2012).

Choosing the bottleneck dimension \(k\)

  • Too small ($k < $ intrinsic dim): the AE underfits — reconstruction error is high even on training data. Phases that differ only subtly get merged into a single latent cluster.
  • Too large ($k > $ intrinsic dim): the AE overfits — extra latent dimensions encode noise rather than signal. Latent clusters spread out and silhouette score drops.
  • Procedure: sweep \(k = 1, 2, \ldots, 2K_{\text{true}}\); plot validation reconstruction error vs \(k\) and look for the elbow. Cross-check with silhouette score — both should peak near the true intrinsic dimension.
  • Sanity check: compare against PCA at the same \(k\). A nonlinear AE should reconstruct at least as well as PCA; if the AE underperforms PCA, it is either under-trained or the bottleneck is too narrow.
  • Rule of thumb: start with $k = $ number of known phases. For an EELS map of an iron-oxide film with 3 known phases, try \(k = 3, 4, 5\); the notebook exercise demonstrates this sweep concretely.

Denoising AE in EM: practical considerations

  • Noise model matters: for Poisson-dominated EELS, use Poisson NLL loss or pre-normalise spectra to roughly constant variance before MSE. A mismatch degrades denoising of low-count bins.
  • Noise level during training: the noise level \(\sigma\) of \(\tilde x_i = x_i + \epsilon_i\) is a hyperparameter. Too low: the AE learns the identity and memorises noise. Too high: the AE cannot reconstruct even clean spectra.
  • Beam-damage awareness: in a real experiment, “noisy” is also “beam-damaged” for longer exposures. The denoising AE must be trained on data with the same noise source as the test data. Mixing Poisson noise with Gaussian readout noise requires a Poisson+Gaussian model.
  • When to prefer PCA denoising: if you have fewer than ~200 spectra, PCA is more robust. The AE needs enough data to learn the spectral manifold. Rule of thumb: AE wins for \(N > 500\) spectra; PCA wins for \(N < 200\).

t-SNE and UMAP: visualising high-dimensional latent codes

Left: t-SNE (perplexity=30) — 2D embedding of 240 latent codes from a 10-D AE. Four iron-oxide phases form islands. Warning: inter-island distances are not metric — the gap between Fe₂O₃ and FeO in this plot does not reflect their spectral similarity. Right: UMAP (n_neighbors=15) — same codes, tighter clusters, better preservation of global structure. Use UMAP for 2026 pipelines; t-SNE is a useful diagnostic.

AE vs VAE: summary

Vanilla AE VAE
Encoder output a point \(\mathbf{z}\) a distribution \(\mathcal{N}(\boldsymbol\mu,\operatorname{diag}\boldsymbol\sigma^2)\)
Loss reconstruction \(-\)ELBO = reconstruction + KL
Latent structure arbitrary scale, holes ≈ \(\mathcal{N}(\mathbf{0},I)\), space-filling
Sampling / interpolation unreliable \(\mathbf{z}\sim\mathcal{N}(\mathbf{0},I)\) → decode
Per-sample uncertainty none \(\boldsymbol\sigma_\phi(\mathbf{x})\) (approximate)
Typical EM use denoising, compression, anomaly scores smooth latent maps, disentangling (rVAE), generation

Generative adversarial networks (GANs) — sharper samples, no encoder, no likelihood — return in Week 13 together with diffusion models as priors for inverse problems Goodfellow, Ian et al., (2014).

rVAE workflow for dynamic STEM data

Workflow of Kalinin, Sergei V et al. (2021) for e-beam-driven dynamics in Si-doped graphene: a deep CNN semantically segments each STEM frame into atom types, sub-images are cropped around atomic positions, and an (r)VAE compresses them into disentangled building blocks (blue: rotation/shift latents; red: content latents). Reproduced from Kalinin et al., Sci. Adv. 7, eabd5084 (2021), open access.

GMM vs AE vs rVAE on the same atomic patches

Descriptors for graphene sub-images with small (top) and large (bottom) windows. (A, D) GMM classes: many statistical components, rotation spread over several classes. (B, E) Convolutional AE latent grid: sublattices separate, but rotation, shift and defects are entangled. (C, F) rVAE latent grid: with rotation removed, only one latent direction varies (degree of crystallinity); with larger windows a second direction encodes proximity to edges. From Kalinin, Sergei V et al. (2021).

Putting it all together: the unsupervised EM pipeline

Complete pipeline from raw EELS/EDS spectra to actionable outputs. Raw spectra enter a denoising AE/VAE encoder; the latent space \(z\) branches into (1) a phase map via k-means/GMM, (2) denoised spectra via the decoder, and (3) anomaly flags where reconstruction error exceeds a threshold. All three outputs require zero labels — and all three need the validation checklist before they become claims.

References

Materials data science, Stefan Sandfeld & others.
Robust design of semi-automated clustering models for 4D-STEM datasets, APL Machine Learning, Alexandra Bruefach, Colin Ophus, & M. C. Scott https://doi.org/10.1063/5.0130546.
Deep learning for multi-dimensional spectral data, npj Computational Materials, Francisco de la Peña & others.
Maximum likelihood from incomplete data via the EM algorithm, Journal of the Royal Statistical Society: Series B (Methodological), Arthur P. Dempster, Nan M. Laird, & Donald B. Rubin.
Pattern recognition and machine learning, Christopher M. Bishop.
Estimating the dimension of a model, The Annals of Statistics, Gideon Schwarz.
Deep learning, Ian Goodfellow, Yoshua Bengio, & Aaron Courville.
Neural networks and principal component analysis: Learning from examples without local minima, Neural Networks, Pierre Baldi & Kurt Hornik.
Extracting and composing robust features with denoising autoencoders, Proceedings of the 25th international conference on machine learning (ICML), Pascal Vincent, Hugo Larochelle, Yoshua Bengio, & Pierre-Antoine Manzagol.
Detection of point defects in STEM images beyond human perception via unsupervised convolutional variational autoencoder, Small, Endrit Prifti, Mathias Kläui, Dierk Raabe, & Benjamin H. Savitzky https://doi.org/10.1002/smll.202303024.
Auto-encoding variational Bayes, International conference on learning representations (ICLR), Diederik P. Kingma & Max Welling.
An introduction to variational autoencoders, Foundations and Trends in Machine Learning, Diederik P. Kingma & Max Welling.
Beta-VAE: Learning basic visual concepts with a constrained variational framework, International conference on learning representations (ICLR), Irina Higgins, Loic Matthey, Arka Pal, Christopher Burgess, Xavier Glorot, Matthew Botvinick, Shakir Mohamed, & Alexander Lerchner.
Explicitly disentangling image content from translation and rotation with spatial-VAE, Advances in neural information processing systems (NeurIPS), Tristan Bepler, Ellen D. Zhong, Kotaro Kelley, Edward Brignole, & Bonnie Berger.
Exploring order parameters and dynamic processes in disordered systems via variational autoencoders, Science Advances, Sergei V Kalinin, Ondrej Dyck, Stephen Jesse, & Maxim Ziatdinov.
How to use t-SNE effectively, Distill, Martin Wattenberg, Fernanda Viégas, & Ian Johnson https://doi.org/10.23915/distill.00002.
Deep learning in electron microscopy, Machine Learning: Science and Technology, Jeffrey M. Ede https://doi.org/10.1088/2632-2153/abd614.
Machine learning: A probabilistic perspective, Kevin P. Murphy.
Generative adversarial nets, Advances in neural information processing systems (NeurIPS), Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, & Yoshua Bengio.