Data Science for Electron Microscopy
Week 13: Inverse problems II: ptychography, generative priors & synthesis

Prof. Dr. Philipp Pelz

FAU Erlangen-Nürnberg

Institute of Micro- and Nanostructure Research

FAU Logo IMN Logo CENEM Logo ERC Logo Eclipse Logo

Recap: Week 12 and today’s question

  • Week 12: every measurement follows \(\mathbf{y} = H\mathbf{x} + \boldsymbol{\epsilon}\); inversion is ill-posed (null space, small singular values), so we added priors (Tikhonov, TV, plug-and-play) and applied them to tomography and sensor fusion.
  • The gap: (1) the forward models were linear — but a detector records \(|\psi|^2\) and loses the phase; (2) the priors were hand-designed.
  • Today:
    1. Ptychography — a non-linear forward model made solvable by redundancy, with the right physics (strong phase, multislice) and the right solver (ePIE or autodiff).
    2. Generative priors — GAN → diffusion, diffusion posterior sampling, and physics constraints as the counterweight.
    3. Limits, causality and trust — when does a powerful prior invent structure?
    4. Course synthesis & exam preparation.

Learning outcomes

By the end of this lecture you can:

  1. Write down the ptychographic forward model and explain why overlap makes phase retrieval over-determined.
  2. State the strong-phase approximation, when it fails, and how multislice fixes it.
  3. Formulate ptychography as loss minimisation and explain what autodiff buys over ePIE.
  4. Explain GANs and diffusion models as learned priors and how diffusion posterior sampling works.
  5. Distinguish soft (loss) and hard (architecture) physics constraints.
  6. Diagnose hallucination and distribution shift before trusting a generative reconstruction.
  7. Place the course’s methods on the E1–E6 explainability ladder.

Road map and self-study

  • Road map (90 min): phase problem & ptychography (11) · object physics: strong phase & multislice (9) · ptychography as gradient descent (11) · learned priors: GAN → diffusion (10) · diffusion posterior sampling (5) · physics constraints (4) · risks, limits & causality (9) · course synthesis (7) · exam & miniproject (4) · wrap-up (8) · questions (8).
  • Self-study: notebooks/week13_ptychography.ipynb — forward model + ePIE with an overlap sweep; the same model reconstructed by autodiff gradient descent; a low-dose exercise with a TV prior in one line of the loss.

The phase problem

  • Detectors record intensity \(|\psi|^2\); the phase \(\phi\) of \(\psi = |\psi|e^{i\phi}\) is discarded.
  • In a thin specimen \(|\psi|\) is nearly uniform — almost all structural information sits in the phase.
  • Recovering \(\phi\) from intensities is the phase problem.

Object phase (left), its Fourier amplitude (centre), and the recorded intensity \(|\mathcal{F}[O]|^2\) (right): the phase is gone.

Redundancy is the key: overlapping probe positions

  • One pattern: \(N^2\) intensities for \(2N^2\) real unknowns — under-determined.
  • Ptychography: scan the probe with a step smaller than the probe; neighbouring positions share most unknowns.
  • 64×64 scan × 64×64 patterns ≈ 16.7 M measurements for a few \(10^4\) object pixels — massively over-determined. Rodenburg, John M. et al., (2007), doi:10.1103/PhysRevLett.98.034801
  • Same redundancy idea as a tomographic tilt series (Week 12), applied to phase.

Probe centres (crosses) and footprints (circles): each object point is illuminated by several positions.

The ptychographic forward model

The four steps at scan position \(j\): (1) crop the object patch \(O_j(\mathbf{r})\); (2) multiply by the probe \(P(\mathbf{r})\) to form the exit wave; (3) FFT to the far field; (4) take \(|\cdot|^2\) to get the measured intensity. Steps 1–3 are reversible; step 4 is not — phase is lost.

\[I_j(\mathbf{k}) = \bigl|\mathcal{F}\bigl[P(\mathbf{r})\cdot O(\mathbf{r}-\mathbf{r}_j)\bigr]\bigr|^2\]

Iterative phase retrieval: ePIE

  • Alternate two constraints: Fourier amplitude must match \(\sqrt{I_j}\) (replace amplitude, keep phase); then propagate the correction back to the object.
  • One ePIE step at position \(j\):
    1. Predict: \(\psi_j = P \cdot O_j\) → \(\Psi_j = \mathcal{F}[\psi_j]\)
    2. Replace amplitude: \(\Psi_j' = \sqrt{I_j} \cdot e^{i\arg \Psi_j}\)
    3. Back-propagate: \(\psi_j' = \mathcal{F}^{-1}[\Psi_j']\)
    4. Update object: \(O_j \leftarrow O_j + \beta \frac{P^*}{|P|^2_{\max}}(\psi_j' - \psi_j)\)
  • Loop over positions in random order until the amplitude-consistency error stops falling. Maiden, Andrew M. et al., (2009), doi:10.1016/j.ultramic.2009.05.012

ePIE convergence: error vs iteration

ePIE amplitude-consistency error vs iteration — output of week13_ptychography.ipynb (SEED=42, 40 iterations). Step 4 px (75 % overlap) → 0.0021; step 6 px (62 %) → 0.0067; step 8 px (50 %) → 0.0091. More overlap, lower error. Low error = self-consistent, not necessarily correct.

The object model: from weak to strong phase

  • Transmission function: \(O(\mathbf{r}) = e^{i\sigma V_p(\mathbf{r})}\), projected potential \(V_p = \int V\,dz\), interaction constant \(\sigma\).
  • Weak-phase approximation (WPOA): \(\sigma V_p \ll 1 \Rightarrow O \approx 1 + i\sigma V_p\) — the forward model becomes linear (basis of CTF theory).
  • Strong-phase approximation (SPA): keep the full exponential, \(\psi_{\text{exit}} = \psi_0\,e^{i\sigma V_p}\) — valid for phase shifts of several radians.
  • SPA assumes the probe does not change shape inside the specimen (no propagation), pure phase object.
  • In the notebook: parameterising \(O = e^{i\phi}\) is the SPA — a hard constraint \(|O|=1\).

When the strong-phase approximation breaks: multislice

  • SPA fails for thick specimens (≳ 10–20 nm), low beam energies (< ~60 keV) and strong scatterers (heavy elements, channelling) — the probe changes shape inside the sample.
  • Symptom: phase no longer scales linearly with thickness; ghost contrast at heavy columns.
  • Multislice Cowley, J. M. et al., (1957), doi:10.1107/S0365110X57002194: cut the specimen into thin slices, transmit and propagate: \[\psi_{n+1} = p(\Delta z) \otimes \bigl[\psi_n \cdot e^{i\sigma V_n}\bigr], \qquad \mathcal{F}[p](\mathbf{k}, \Delta z) = e^{-i\pi\lambda\Delta z |\mathbf{k}|^2}\]
  • Multislice ptychography reconstructs all slices \(V_1\ldots V_N\) Maiden, Andrew M. et al., (2012), doi:10.1364/JOSAA.29.001606 — a deep, differentiable chain of multiply → FFT → propagate → IFFT.

Phase object vs multislice: a picture

Left: the focused probe changes shape over tens of nm of defocus — comparable to the sample thickness (dashed box). Centre: the phase-object (SPA) model treats the whole specimen as one phase screen, \(\psi_{\text{exit}} = \psi_{\text{probe}}\,e^{i\sigma V}\). Right: multislice theory alternates transmission through thin slices \(V_n\) and Fresnel propagation \(p(\Delta z)\). Slide credit: D. Muller (M&M 2023), after Chen, Zhen et al. (2021), doi:10.1126/science.abg2533.

Depth sectioning, its missing wedge, and the thermal limit

  • Depth information is carried only by high transverse frequencies (propagator phase \(\pi\lambda\Delta z|\mathbf{k}|^2\)).
  • Low frequencies are nearly depth-invariant → a missing wedge in \(k_z\), as in limited-angle tomography (Week 12): depth resolution a few nm vs < 0.5 Å laterally.
  • Resolution limit: in PrScO₃ multislice ptychography reaches ~20 pm — set by thermal vibrations of the atoms, not the microscope. Chen, Zhen et al., (2021), doi:10.1126/science.abg2533

Multislice ptychography of PrScO₃: experiment matches simulation only once thermal vibrations are included. After Chen, Zhen et al. (2021), doi:10.1126/science.abg2533; slide credit D. Muller.

Ptychography as loss minimisation

  • Loss over all scan positions: \[\hat{\phi} = \arg\min_{\phi} \; \sum_j \Bigl\|\, \bigl|\mathcal{F}[P\cdot O_j(\phi)]\bigr| - \sqrt{I_j}\,\Bigr\|_2^2 \;+\; \lambda\,R(\phi)\]
  • Data term from the noise model (Week 2): least squares on \(I\) (Gaussian), Poisson NLL (counts), or the amplitude loss on \(\sqrt{I}\) ≈ variance-stabilised Poisson — the common choice.
  • Parameters: object, and optionally probe, scan positions, slice potentials, tilt angles.
  • Solver: any gradient method (Week 4). ePIE ≈ SGD with batch size 1 and a hand-derived step \(\beta/|P|^2_{\max}\).

Automatic differentiation: write the forward model, get the gradient

def forward(phi, probe, rows, cols):
    O = torch.exp(1j * phi)                    # SPA: pure phase object (hard constraint |O| = 1)
    exit_waves = probe * O[rows, cols]         # probe x all object patches at once
    return torch.abs(torch.fft.fft2(exit_waves))   # predicted diffraction amplitudes

phi = torch.zeros(N, N, requires_grad=True)
opt = torch.optim.Adam([phi], lr=0.02)
for it in range(200):
    loss = ((forward(phi, probe, rows, cols) - meas_amps) ** 2).sum()   # data fidelity
    loss = loss + lam_tv * tv(phi)                                      # prior: one extra line
    opt.zero_grad(); loss.backward(); opt.step()                        # autograd does the calculus

ePIE vs autodiff gradient descent: same problem, two solvers

Amplitude-consistency error per pass through all 81 diffraction patterns (noise-free, 75 % overlap). Output of week13_ptychography.ipynb.
after 200 passes ePIE autodiff GD
amplitude error \(7.7\times10^{-4}\) \(8.7\times10^{-5}\)
aligned phase RMS \(1.7\times10^{-3}\) rad \(2.9\times10^{-4}\) rad
wall time (CPU) 4.5 s 9.3 s
  • ePIE is fast early; full-batch Adam is more precise late (10× lower error).

What automatic differentiation buys

  • Flexibility: probe refinement, position correction, mixed probe states, multislice, tilt series — each is one more parameter or layer.
  • Any loss, any optimiser: Poisson NLL at low dose; regularisers as loss terms; mini-batch SGD over millions of patterns on GPUs.
  • Composability: a denoiser, a generative prior or a physics residual plugs into the same graph — the bridge to the rest of today.
  • Costs: memory, hyperparameters, non-convexity — initialisation still matters.

Low dose: a prior in one line of the loss

  • Notebook exercise 2: Poisson noise, 100 e⁻ per pattern; aligned phase RMS error:
method phase RMS [rad]
ePIE, 40 iterations 0.094
autodiff GD, \(\lambda_{TV}=0\) (200 it.) 0.103
autodiff GD, \(\lambda_{TV}=1\) 0.080
autodiff GD, \(\lambda_{TV}=8\) 0.140
  • U-curve: too little prior fits noise; too much flattens the phase (bias) — the Week 12 \(\lambda\) trade-off.
  • Early stopping is an implicit regulariser: ePIE at 40 iterations beats unregularised GD.
  • The weaker the data, the more the prior has to carry.

From hand-built regularisers to learned priors

  • Regularisation = MAP: \[\min_{\mathbf{x}} \|H(\mathbf{x}) - \mathbf{y}\|^2 + \lambda R(\mathbf{x}) \;\Leftrightarrow\; p(\mathbf{x}) \propto e^{-\lambda R(\mathbf{x})}\]
  • Tikhonov, TV, positivity: hand-designed; none knows what real specimens look like.
  • Two ways to do better: put known physics into the loss/architecture, or learn \(p(\mathbf{x})\) from data with a generative model.
  • Week 12’s plug-and-play denoiser is the bridge: a diffusion model is a denoiser at every noise level.

Hand-designed prior: a simple set (left). Learned prior: an irregular manifold shaped by the training data (right) — accurate in distribution, dangerous outside it.

GANs: an adversarial prior for sharp images

  • Minimax game Goodfellow, Ian et al., (2014): \[\min_G \max_D \; \mathbb{E}_{\mathbf{x}\sim p_{\text{data}}}[\log D(\mathbf{x})] + \mathbb{E}_{\mathbf{z}}[\log(1 - D(G(\mathbf{z})))]\] generator \(G(\mathbf{z})\) maps noise to an image; discriminator \(D(\mathbf{x})\) says “real or fake”.
  • Strength: sharp samples in one forward pass (no MSE blur); conditional GANs map low-dose → high-dose.
  • As a prior: \(\hat{\mathbf{z}} = \arg\min_{\mathbf{z}} \|H(G(\mathbf{z})) - \mathbf{y}\|^2\) — the solution lies on the generator’s manifold.
  • Failure modes: unstable training, mode collapse, and no likelihood.

From GANs to diffusion: why the field moved

  • Stable training: a plain regression loss (predict the noise) — one network, no adversary.
  • Mode coverage: likelihood-based objective → diverse samples, no mode collapse.
  • A score, not just a sampler: diffusion gives \(\nabla_{\mathbf{x}} \log p_t(\mathbf{x})\) at every noise level — what Bayesian inverse problems need. Song, Yang et al., (2021)
  • Price: slow sampling — tens to hundreds of network evaluations vs one for a GAN.

Diffusion models: forward noising, learned reverse

  • Forward process (fixed): add Gaussian noise with schedule \(\beta_t\); closed-form marginal Ho, Jonathan et al., (2020) \[q(\mathbf{x}_t \mid \mathbf{x}_0) = \mathcal{N}\bigl(\sqrt{\bar\alpha_t}\,\mathbf{x}_0,\,(1-\bar\alpha_t)\mathbf{I}\bigr), \qquad \bar\alpha_t = \textstyle\prod_{s\le t}(1-\beta_s)\]
  • Training: form \(\mathbf{x}_t = \sqrt{\bar\alpha_t}\mathbf{x}_0 + \sqrt{1-\bar\alpha_t}\,\boldsymbol\epsilon\) and minimise \[\mathcal{L} = \mathbb{E}\,\bigl\|\boldsymbol\epsilon - \boldsymbol\epsilon_\theta(\mathbf{x}_t, t)\bigr\|^2\] with a U-Net (Week 7) that predicts the noise.
  • Sampling: start from \(\mathbf{x}_T\sim\mathcal{N}(0,I)\), denoise step by step — coarse structure first.
  • Score: \(\nabla_{\mathbf{x}_t}\log p_t(\mathbf{x}_t) \approx -\boldsymbol\epsilon_\theta(\mathbf{x}_t,t)/\sqrt{1-\bar\alpha_t}\).

Diffusion posterior sampling (DPS)

  • Bayes for scores — the normaliser drops out: \[\nabla_{\mathbf{x}_t}\log p_t(\mathbf{x}_t\mid\mathbf{y}) = \underbrace{\nabla_{\mathbf{x}_t}\log p_t(\mathbf{x}_t)}_{\text{learned prior score}} + \underbrace{\nabla_{\mathbf{x}_t}\log p_t(\mathbf{y}\mid\mathbf{x}_t)}_{\text{data consistency (intractable)}}\]
  • DPS Chung, Hyungjin et al., (2023): evaluate the likelihood at the clean-image (Tweedie) estimate \(\hat{\mathbf{x}}_0 = (\mathbf{x}_t - \sqrt{1-\bar\alpha_t}\,\boldsymbol\epsilon_\theta)/\sqrt{\bar\alpha_t}\).
  • Each reverse step: usual denoising step → \(\mathbf{x}'_{t-1}\), then \[\mathbf{x}_{t-1} = \mathbf{x}'_{t-1} - \zeta_t\,\nabla_{\mathbf{x}_t}\bigl\|\mathbf{y} - H(\hat{\mathbf{x}}_0)\bigr\|\]
  • Works for non-linear \(H\) (ptychography): autodiff through network and physics. Repeat with new seeds → an ensemble of posterior samples.

DPS on a toy: what the data fix and what the prior decides

2-D toy with an exact prior (four Gaussian “structural variants”) and exact noised scores. We observe only \(x_1\) (a projection, \(H=[1\;0]\)); \(x_2\) lies in the null space of \(H\). Left: unconditional reverse diffusion reproduces the prior. Centre: exact posterior. Right: DPS samples. Generated by img/make_figures.py.

Physics-informed learning: the key idea

  • Generative priors know data, not physics. Physics-informed learning adds a known equation, conservation law or symmetry to the objective. Raissi, Maziar et al., (2019), doi:10.1016/j.jcp.2018.10.045
  • PINN loss: \[\mathcal{L}(\theta) = \underbrace{\tfrac{1}{N}\textstyle\sum_i \|f_\theta(\mathbf{x}_i) - y_i\|^2}_{\text{data fidelity}} + \lambda\,\underbrace{\tfrac{1}{M}\textstyle\sum_m \|\mathcal{N}[f_\theta](\mathbf{x}_m)\|^2}_{\text{physics residual}}\]
  • Small-data advantage: the residual constrains the solution where there are no data.

Soft vs hard physics constraints

  • Soft (penalty): add \(\lambda\|\text{violation}\|^2\) to the loss — flexible, only approximately satisfied, λ must be tuned.
  • Hard (parameterisation / architecture): make violations impossible.
    • \(O = e^{i\phi}\) guarantees \(|O|=1\) (notebook); positivity via softplus.
    • Boundary conditions — Lagaris substitution Lagaris, Isaac G. et al., (1998): \(\hat u(x) = A(x) + B(x)\,N_\theta(x)\).
    • Symmetry: equivariant architectures (CNN, GNN; Weeks 7, 10).
  • Hard constraints are exact but only as good as their physics: wrong physics (SPA for a thick sample) biases the result confidently.

The honest risks: hallucination in EM

Low-dose HAADF image (centre) of four real atomic columns. A generative denoiser (right) recovers them but also invents a fifth (red arrow) — no basis in the ground truth (left). The prior places an atom where one “usually is”.

Distribution shift: when training data do not match the specimen

  • Prior learned from \(p_{\text{train}}\), specimen from \(p_{\text{test}} \neq p_{\text{train}}\) — the most common cause of hallucination.
  • EM examples: prior trained on perfect SrTiO₃ smooths away a grain boundary; Au-trained prior gives Pt columns Au-like contrast; simulation-trained prior on experimental data.
  • Detect: residual \(\mathbf{y} - H(\hat{\mathbf{x}})\) should look like noise; OOD scores (Week 11); disagreement with a prior-free (ePIE / TV) reconstruction.
  • Mitigate: diverse training data, physical forward model in the loop (DPS, not image-to-image), independent validation.

When NOT to trust a generative reconstruction

  • Red flags:
    1. Very low dose and a suspiciously clean, noise-free result.
    2. A feature of a type absent from the training data.
    3. Structured residuals.
    4. The feature would be the headline result of your paper.
  • Checks:
    1. Ensemble agreement across posterior samples / seeds.
    2. Normalised residuals \(r_i = (y_i - H(\hat{\mathbf{x}})_i)/\sqrt{H(\hat{\mathbf{x}})_i} \sim \mathcal{N}(0,1)\).
    3. Agreement with a prior-free (ePIE / TV) reconstruction.
    4. Key features confirmed at higher dose or independently.

Limits: more powerful = more assumptions

  • No algorithm recovers what the measurement never contained: null space of \(H\) (missing wedge, lost phase), frequencies below the noise floor, structure below the thermal blur.
  • A prior fills it with plausible content. Plausible ≠ measured.
  • Hierarchy: overlap alone (the forward model is right) → + TV (piecewise smooth) → + physics residual (the equation is right) → + generative prior (looks like the training set).
  • Each step buys quality at low dose and costs verifiability: state, test and report the assumptions.

Causality vs correlation in EM analyses

  • ML finds associations; materials science needs mechanisms. A defect classifier trained on furnace A may learn “furnace A ↔︎ defects” and fail on furnace B.
  • Reconstruction version: a generative prior “explains” the data with the most typical structure — a statistical, not a causal, argument.
  • Only interventions establish causes: inject an artefact (Week 7 shortcut demo), re-image at higher dose, measure a fresh specimen.
  • Honest validation (Week 4): split by specimen/session to learn the physics, not the specimen.

The 13-week arc as one methodology

The WS 26/27 course arc: data, noise and honest learning (W1–4); representations and model families from trees to transformers (W5–10); knowing what we don’t know (W11); physics-based reconstruction (W12–13). Cross-cutting threads connect them.

The E1–E6 explainability ladder — collected over 13 weeks

The six levels of explainability Neuer, Michael et al., (2024), ordered from data to decision, with the weeks in which each level was taught. Explainability was not a separate topic: every model family came with its own explanation tool.

Recurring trust failures across the course

The four recurring trust failures: leakage (W4), shortcut learning (W7), miscalibration and out-of-distribution inputs (W11), hallucination (W13). All share one root cause: a statistical association or prior replaced the physical signal, and a single accuracy number could not reveal it.

Five pillars of trustworthy EM data science

  1. Correct physics: the noise model decides the loss, the forward model decides the inversion.
  2. Honest validation: no leakage; group splits; report generalisation.
  3. Calibrated confidence: 95 % intervals that contain the truth 95 % of the time — checked.
  4. Faithful explanation: attributions verified by sanity checks and interventions.
  5. Expert validation: a materials scientist checks that it makes physical sense.
  • Collect the right data, build the right model, know what you don’t know, and be able to explain why you trust the answer.

Exam preparation

  • Written exam: 60 % of the grade (miniproject 40 %); format and date as announced on the course page.
  • Question types:
    • Conceptual — “When does the strong-phase approximation fail?”
    • Application — “A saliency map highlights a scale bar. What do you conclude, and which experiment confirms it?”
    • Calculation — noise model → loss, PCA/SVD, ridge/Tikhonov, backprop, GP posterior, Radon, ePIE update.
    • Judgement — “A colleague denoises a 50 e⁻/Ų image with a diffusion model trained on simulations. Three risks, one check each.”
  • Study material: _shared/exam_mustknow.md. For every week: core idea in one sentence, key equation, one EM application and one failure mode.

Miniproject: final checklist

  • Specification: _shared/miniproject.md — options A–E.
  • Rubric (40 %): framing 15 · methodology 25 · uncertainty & validation 20 · explainability 15 · reproducibility 15 · report 10.
  • Explainability with a materials conclusion; for Option D show what the prior contributes (with/without regularisation, residual plot).
  • Reproducibility: jupyter nbconvert --to notebook --execute your_notebook.ipynb must run end-to-end.

Self-study notebook

  • Notebook: notebooks/week13_ptychography.ipynb — Open In Colab
  • Parts 1–5: forward model → ePIE (0.0991 → 0.0021); overlap sweep 4 / 6 / 8 px → 0.0021 / 0.0067 / 0.0091.
  • Part 6: autodiff gradient descent with \(O=e^{i\phi}\) — \(8.7\times10^{-5}\) vs \(7.7\times10^{-4}\) (ePIE) after 200 passes.
  • Part 7: 100 e⁻/pattern, TV prior in one line; find the best \(\lambda_{TV}\) and explain the U-curve.
  • CPU only, < 1 min.

Summary

  • Ptychography: overlap makes phase retrieval over-determined; ePIE solves it by alternating projections.
  • Object physics: weak phase → strong phase → multislice for thick samples; depth has a missing wedge; thermal vibrations set the final limit.
  • Autodiff gradient descent solves the same loss with any parameters, losses and priors.
  • Generative priors: GANs → diffusion (stable, diverse, gives a score); DPS adds a data-consistency gradient through Tweedie’s estimate.
  • Physics constraints: soft (loss) vs hard (parameterisation); wrong physics biases confidently.
  • Trust: low dose = strong prior = hallucination risk — residuals, ensembles, prior-free baselines.
  • Synthesis: E1–E6 ladder and five pillars — more powerful = more assumptions.

Must-know takeaways (Week 13)

  1. Detectors record \(|\psi|^2\); ptychographic overlap makes phase retrieval over-determined: \(I_j = |\mathcal{F}[P\cdot O(\mathbf{r}-\mathbf{r}_j)]|^2\).
  2. ePIE: replace the Fourier amplitude with \(\sqrt{I_j}\), back-transform, update \(O_j \mathrel{+}= \beta\,P^*/|P|^2_{\max}\,(\psi'_j-\psi_j)\).
  3. SPA: \(\psi_{\text{exit}} = \psi_0\,e^{i\sigma V_p}\); fails for thick samples / strong scatterers → multislice (transmit, propagate, repeat).
  4. Ptychography = loss minimisation; autodiff differentiates through FFT and modulus, so probe, positions, slices and priors are just extra terms.
  5. Posterior score = prior score + likelihood score; DPS approximates the likelihood at the Tweedie estimate \(\hat{\mathbf{x}}_0\).
  6. Hard constraints (parameterisation) are exact but only as good as their physics; soft constraints need a λ.
  7. The weaker the data, the more the prior decides — check residuals, ensembles and a prior-free baseline before trusting a generative reconstruction.

Closing

  • Over 13 weeks you built a data-science toolkit for EM — from noise models, trees, CNNs, transformers and GPs to physics-based and generative reconstruction.
  • The scientific obligation: you are responsible for validating every number, explanation and reconstruction against physics.
  • Next steps: exam_mustknow.md, miniproject.md, and the Ai4Mat companion notebooks.
  • Thank you for 13 weeks of work.

Continue

Backup slides

Material that did not fit the 90-minute lecture path — for self-study and questions.

Beyond one view: ptychographic tomography

  • Combine Week 12 and Week 13: reconstruct a phase projection (or multislice stack) at every tilt angle, then solve a tomography problem — or do both at once.
  • Ptychographic atomic electron tomography: 34.5 million diffraction patterns → atomic-resolution tilt series of a double-wall carbon nanotube with an encapsulated ZrTe structure; revealed a previously unobserved ZrTe₂ phase. Pelz, Philipp M. et al., (2023), doi:10.1038/s41467-023-43634-z
  • Multislice ptychographic tomography: 2 Å axial and 0.7 Å lateral resolution in a volume beyond the depth-of-field limit — a 13.5× axial improvement over single-view multislice. Romanov, Andrey et al., (2024), doi:10.1088/2515-7639/ad9ad2
  • End-to-end: reconstruct the 3-D potential directly from the unaligned 4D-STEM tilt series, including alignment and a restricted tilt range — sub-Å near-isotropic resolution. You, Shengbo et al., (2024), doi:10.1088/1402-4896/ad9a1a
  • Why this became possible: all of these are one big differentiable forward model optimised by gradient descent.

What ptychography buys: dose–resolution trade-off

Schematic dose–resolution trade-off for ADF-STEM (red) and ptychography (blue); y-axis inverted (higher = finer). In the low-dose regime both scale as \(d \propto 1/\sqrt{\text{dose}}\), but ptychography uses all scattered electrons and achieves better resolution at equal dose. At high dose ADF saturates at the probe-size limit; ptychography keeps improving because it deconvolves the probe. Chen, Zhen et al., (2021), doi:10.1126/science.abg2533

Experimental parameters and quality control

  • What you choose: probe size / defocus (aberration corrector settings), scan step (~20–40 % of the probe diameter), camera length (detector angular range ↔︎ real-space pixel size), total dose.
  • Detector sampling: the diffraction pattern must be Nyquist-sampled; under-sampling aliases (Week 2).
  • Partial coherence: model as mixed probe states, \(I_j = \sum_k |\mathcal{F}[P_k \cdot O_j]|^2\) — trivial to add in an autodiff framework.
  • What you get: complex object \(O(\mathbf{r})\) (amplitude + phase), reconstructed probe(s), refined positions.
  • Quality check: amplitude error decreasing smoothly and plateauing well below its start; residual consistent with Poisson noise; reconstruction stable across initialisations and data halves (a half-set comparison, like cryo-EM FSC). Maiden, Andrew M. et al., (2009), doi:10.1016/j.ultramic.2009.05.012

The plug-and-play framework: any denoiser as a prior

  • Insight: the proximal step of a regulariser \(R\) has the same form as Gaussian denoising. Kamilov, Ulugbek S. et al., (2023), doi:10.1109/MSP.2022.3199595
  • Plug-and-play (PnP): replace the proximal step in ADMM / proximal gradient descent with any strong denoiser — BM3D, a DnCNN or a U-Net (Week 7) trained on EM images.
  • Algorithm: alternate
    1. Data step: \(\mathbf{x} \leftarrow \mathbf{x} - \eta \nabla_{\mathbf{x}} \|H(\mathbf{x}) - \mathbf{y}\|^2\)
    2. Prior step: \(\mathbf{x} \leftarrow D_\sigma(\mathbf{x})\) (apply the learned denoiser)
  • Result: the denoiser defines the prior implicitly. Diffusion models are the logical extreme: a denoiser trained at every noise level.

Tweedie’s formula: a denoiser gives a clean-image estimate

  • At any step \(t\) the network gives an estimate of the clean image (Tweedie): \[\hat{\mathbf{x}}_0(\mathbf{x}_t) = \frac{\mathbf{x}_t + (1-\bar\alpha_t)\,\nabla_{\mathbf{x}_t}\log p_t(\mathbf{x}_t)}{\sqrt{\bar\alpha_t}} = \frac{\mathbf{x}_t - \sqrt{1-\bar\alpha_t}\,\boldsymbol\epsilon_\theta(\mathbf{x}_t,t)}{\sqrt{\bar\alpha_t}}\]
  • \(\hat{\mathbf{x}}_0\) is the posterior mean \(\mathbb{E}[\mathbf{x}_0 \mid \mathbf{x}_t]\) — blurry early (large \(t\)), sharp late.
  • Why this matters for inverse problems: we can push \(\hat{\mathbf{x}}_0\) through the physical forward model \(H\) at every step and compare with the measurement \(\mathbf{y}\) — even though \(\mathbf{x}_t\) itself is noisy.
  • This is the one ingredient that turns an unconditional generator into a reconstruction algorithm.

Generative models for EM: applications

  • Denoising at low dose: learned priors preserve atomic contrast far better than BM3D/NLM at extreme noise; self-supervised training (Noise2Void / blind-spot) avoids paired clean data.
  • Super-resolution & inpainting: fill in sub-sampled scans (dose reduction by sparse scanning, compressed sensing) — conditioned on measured pixels.
  • Reconstruction priors: diffusion/PnP priors inside ptychography and tomography — especially to fill the missing wedge.
  • Synthetic training data & inverse design: generate realistic micrographs for training segmentation (Weeks 7–8) and crystal/microstructure candidates (MG Unit 12).
  • The common question in every case: did the network reveal structure or invent it? Low dose = weak likelihood = strong prior = high risk.

Choosing a reconstruction method

Decision table: classical regularisation (Tikhonov/TV) for well-understood physics and limited data; learned priors (GAN/diffusion) when a large dataset of similar specimens exists; physics-informed learning when governing equations are known and data are scarce. Combinations (physics + learned prior, e.g. DPS with a physical forward model) are the active research direction.

The explainability & trust toolkit at a glance

Method Question it answers Works on Week
Coefficients / partial dependence How does the output change with a feature? Linear / additive models W4
Permutation importance, SHAP Lundberg, Scott M. et al., (2017) Which features matter, globally and per sample? Any model W5
Saliency, Grad-CAM Selvaraju, Ramprasaath R. et al., (2017), occlusion Which pixels drove this prediction? CNNs (occlusion: any) W7
Latent-space maps, t-SNE/UMAP What structure did the model find? (beware distances) AEs / VAEs / embeddings W9
Attention maps Where did the transformer look? (≠ why) Transformers W10
Calibration, conformal, OOD gate How sure — and is this input in-distribution? Any predictor W11
Residuals, posterior ensembles, half-sets Is this reconstruction supported by the data? Inverse problems W12–13
  • Sanity check every explanation — e.g. randomise the weights: if the saliency map does not change, it explains the image, not the model. Adebayo, Julius et al., (2018)

References

Hard-x-ray lensless imaging of extended objects, Physical Review Letters, John M. Rodenburg, A. C. Hurst, A. G. Cullis, B. R. Dobson, F. Pfeiffer, O. Bunk, C. David, K. Jefimovs, & I. Johnson https://doi.org/10.1103/PhysRevLett.98.034801.
An improved ptychographical phase retrieval algorithm for diffractive imaging, Ultramicroscopy, Andrew M. Maiden & John M. Rodenburg https://doi.org/10.1016/j.ultramic.2009.05.012.
The scattering of electrons by atoms and crystals. I. A new theoretical approach, Acta Crystallographica, J. M. Cowley & A. F. Moodie https://doi.org/10.1107/S0365110X57002194.
Ptychographic transmission microscopy in three dimensions using a multi-slice approach, Journal of the Optical Society of America A, Andrew M. Maiden, Martin J. Humphry, & John M. Rodenburg https://doi.org/10.1364/JOSAA.29.001606.
Electron ptychography achieves atomic-resolution limits set by lattice vibrations, Science, Zhen Chen, Michal Odstrcil, Yi Jiang, Yimo Han, Ming-Hui Chiu, Lain-Jong Li, & David A. Muller https://doi.org/10.1126/science.abg2533.
Using automatic differentiation as a general framework for ptychographic reconstruction, Optics Express, Saugat Kandel, Siddharth Maddali, Marc Allain, Stephan O. Hruszkewycz, Chris Jacobsen, & Youssef S. G. Nashed https://doi.org/10.1364/OE.27.018653.
A plug-and-play image reconstruction framework, IEEE Signal Processing Magazine, Ulugbek S. Kamilov, Charles A. Bouman, Gregery T. Buzzard, & Brendt Wohlberg https://doi.org/10.1109/MSP.2022.3199595.
Generative adversarial nets, Advances in neural information processing systems, Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, & Yoshua Bengio.
Score-based generative modeling through stochastic differential equations, International conference on learning representations (ICLR), Yang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar, Stefano Ermon, & Ben Poole.
Denoising diffusion probabilistic models, Advances in neural information processing systems, Jonathan Ho, Ajay Jain, & Pieter Abbeel.
Diffusion posterior sampling for general noisy inverse problems, International conference on learning representations (ICLR), Hyungjin Chung, Jeongsol Kim, Michael T. Mccann, Marc L. Klasky, & Jong Chul Ye.
Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations, Journal of Computational Physics, Maziar Raissi, Paris Perdikaris, & George E. Karniadakis https://doi.org/10.1016/j.jcp.2018.10.045.
Artificial neural networks for solving ordinary and partial differential equations, IEEE Transactions on Neural Networks, Isaac G. Lagaris, Aristidis Likas, & Dimitrios I. Fotiadis.
Solving complex nanostructures with ptychographic atomic electron tomography, Nature Communications, Philipp M. Pelz, Sinéad M. Griffin, Scott Stonemeyer, Derek Popple, Hannah DeVyldere, Peter Ercius, Alex Zettl, Mary C. Scott, & Colin Ophus https://doi.org/10.1038/s41467-023-43634-z.
Machine learning for engineers: Introduction to physics-informed, explainable learning methods for AI in engineering applications, Michael Neuer & others.
Multi-slice electron ptychographic tomography for three-dimensional phase-contrast microscopy beyond the depth of focus limits, Journal of Physics: Materials, Andrey Romanov, Min Gee Cho, Mary Cooper Scott, & Philipp Pelz https://doi.org/10.1088/2515-7639/ad9ad2.
Near-isotropic sub-Ångstrom 3d resolution phase contrast imaging achieved by end-to-end ptychographic electron tomography, Physica Scripta, Shengbo You, Andrey Romanov, & Philipp M. Pelz https://doi.org/10.1088/1402-4896/ad9a1a.
A unified approach to interpreting model predictions, Advances in neural information processing systems, Scott M. Lundberg & Su-In Lee.
Grad-CAM: Visual explanations from deep networks via gradient-based localization, Proceedings of the IEEE international conference on computer vision, Ramprasaath R. Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, & Dhruv Batra.
Sanity checks for saliency maps, Advances in neural information processing systems (NeurIPS), Julius Adebayo, Justin Gilmer, Michael Muelly, Ian Goodfellow, Moritz Hardt, & Been Kim.