Skip to content

Instantly share code, notes, and snippets.

@regstuff
Last active September 3, 2026 07:00
Show Gist options
  • Select an option

  • Save regstuff/d68a31ee7aec3ee0aff5e6959fb84b09 to your computer and use it in GitHub Desktop.

Select an option

Save regstuff/d68a31ee7aec3ee0aff5e6959fb84b09 to your computer and use it in GitHub Desktop.
Blackholes, Information & Entropy

Entropy

Entropy in a quantum information context can be said to represent the amount of uncertainty in our knowledge of a system. A completely mixed state has maximum uncertainty in the sense that it is a complete classical ensemble of multiple possibilities that are equally possible, and thus we can say nothing about which outcome is likely. A pure state however is something that we have absolute knowledge of. We know that it is in that state. Note that a pure state can be a superposition. This does not contribute to its uncertainty in terms of entropy. A superposition and its probability amplitudes can change depending on the measurement basis we choose. That is a manifestation of quantum randomness and not uncertainty in the state of the system.

Von Neumann Entropy

The Von Neumann entropy of a system with a density matrix $\rho$ is defined as,

$$S(\rho)=-\text{Tr}(\rho\log_2\rho)$$

where the logarithm is taken in base 2. The logarithm of a matrix is constructed by taking the log of the eigenvalues and rebuilding the resultant matrix with the same original eigenstates via spectral decomposition.

If we diagonalize $\rho$, we end up with the eigenvalues i.e. the probabilities of the ensemble represented by $\rho$, along the diagonal of the matrix. Thus, taking the trace is just the sum of the eigenvalues. Therefore,

$$S=-\sum_ip_i\log_2p_i$$

which is just the Shannon Entropy. When the probability is 0, we assume that $p\log p$ equates to 0 since $x$ tends to 0 faster than $\log x$ tends to negative infinity as $x\to0$.

Entropy Relations

There are a few mathematical equations and inequalities that pertain to the entropy of a single system state as well as multipartite states. Note that in terms of notation, when the context is clear, $S(\rho_A)$ and $S(\rho_B)$ are just written as $A$ and $B$ or $S_A$ and $S_B$.


$$0\le S\le\log_2d$$

where $d$ is the dimension of the Hilbert space of the state. $S=0$ when the state is pure, and the maximum is reached when the state is maximally mixed i.e. there are $d$ states in the ensemble each with a probability of $1/d$. This can be related to the uncertainties in the two situations as explained above. Thus entropy is a measure of the purity of the state. The lower it is, the closer it is to purity.


If $\vert{}\psi\rangle_{AB}$ is a pure state and $\rho_A$ and $\rho_B$ are the two reduced density matrices of this state, then $S(\rho_A)=S(\rho_B)$. This makes sense since if $AB$ is a pure state, then there is no uncertainty in its state. Any uncertainty in $A$ must be "clarified" or negated by information in $B$ and vice versa, so the amount of uncertainties in $A$ and $B$ must be equal. This relationship can be proved via the Schmidt decomposition.

From the above two points, it follows that if $AB$ is a product state of two pure states $A$ and $B$, then $S_A=S_B=0$. If $A$ and $B$ are entangled, then $S_{AB}$ may be zero (if it is pure) or non-zero (if it is mixed), but $S_A$ and $S_B$ are definitely non-zero because they will be mixed. Thus, entropy in a bipartite system is a measure of entanglement. Also note that $S_A$ and $S_B$ need not be equal when the bipartite system is mixed.


If $U$ is a unitary operator, $S(U\rho U^\dagger)=S(\rho)$.


The relative entropy between two states $\rho$ and $\sigma$ is defined as

$$D(\rho\vert{}\vert{}\sigma)=\text{Tr}(\rho\log\rho-\rho\log\sigma)$$

Note that it is assumed here that $\text{supp}(\rho)\subseteq\text{supp}(\sigma)$. This essentially means that $\sigma$ has no zero eigenvalues for any non-zero eigenvalue state of $\rho$; however, $\sigma$ is permitted to have non-zero eigenvalues even where $\rho$ has zero eigenvalues. This is essential if the entropy expression should not blow up to infinity.

Relative entropy is bounded by the trace norm (via Pinsker's inequality),

$$D(\rho\vert{}\vert{}\sigma)\ge\frac{1}{2\ln2}\vert{}\vert{}\rho-\sigma\vert{}\vert{}_1^2$$

where the trace norm of an operator is $\vert{}\vert{}O\vert{}\vert{}_1=\text{Tr}(\sqrt{O^\dagger O})$.

Relative entropy therefore has to be greater than or equal to distinguishability, and it therefore dictates the upper bound on how distinguishable two states are.


$$S_A+S_B\ge S_{AB}$$

is called the subadditivity rule, with the equality holding only if $AB$ is a product state. A product state of pure states will set both sides of the equality to 0. A product state of mixed states will make them non-zero.

The physical intuition behind the inequality is that for an entangled system, information is entangled across both subsystems. If you examine just one or the other system in isolation, you would expect to be more uncertain about the overall system state than if you could inspect the entire bipartite system.

To see how entanglement physically manifests in this inequality, track the evolution of two separate qubits initially prepared in pure states. Before interaction, the system is a strict tensor product. $S(\rho_A)=0$, $S(\rho_B)=0$, and the joint entropy $S(\rho_{AB})=0$. Apply a unitary entangling gate, such as a CNOT, to the physical qubits. Because unitary operations preserve purity, the global state remains perfectly pure. The joint entropy $S(\rho_{AB})$ remains exactly 0. However, the subsystems are now entangled, and the reduced density matrices are mixed. $S(\rho_A)>0$ and $S(\rho_B)>0$. The subadditivity inequality is now strictly enforced.

In classical statistical mechanics, joint entropy can never be less than the entropy of an individual subsystem ($S_{AB}\ge S_A$). The whole must contain at least as much uncertainty as any single part.

Quantum entanglement fundamentally violates this classical bound. Because a maximally entangled state yields $S(\rho_{AB})=0$ while $S(\rho_A)>0$, quantum mechanics permits a total system to possess strictly less entropy than its constituent parts. You can possess absolute, perfect mathematical knowledge of the global composite state while remaining completely ignorant of the local, physical subsystems.


The Araki-Lieb inequality states:

$$\vert{}S_A-S_B\vert{}\le S_{AB}$$

To understand the physical intuition, note that entanglement is a mechanism to reduce uncertainty in individual systems by building correlations across the bipartite system. However, if system $A$ has only 2 units worth of information to entangle compared to 10 units of information in $B$, we would be left with a net uncertainty in the bipartite system of 8 units. This is what the Araki-Lieb inequality dictates: the joint uncertainty $S(AB)$ must always be greater than or equal to the uncompensated, leftover entropy.


From the above two rules, we have bounds on the entropy of a bipartite system:

$$\vert{}S_A-S_B\vert{}\le S_{AB} \le S_A+S_B$$


The Strong subadditivity (SSA) relation states that:

$$S_{AB}+S_{BC}\ge S_{ABC}+S_B$$

To understand what this means intuitively and physically, the inequality is best translated into two equivalent principles: "Information never hurts" and the "Monogamy of Entanglement."

Information Never Hurts (Conditional Entropy): In information theory, conditional entropy $S(A\vert{}B)=S(AB)-S(B)$ measures how much uncertainty remains about subsystem $A$ assuming you already have full access to subsystem $B$.

By simply rearranging the terms in the SSA inequality, we obtain:

$$S(ABC)-S(BC)\le S(AB)-S(B)$$

$$S(A\vert{}BC)\le S(A\vert{}B)$$

Having access to a larger portion of the universe can never make you more ignorant about a specific system. If you are trying to determine the state of $A$, having access to both $B$ and $C$ guarantees you will be left with equal or less uncertainty than if you only had access to $B$. The extra information provided by $C$ might help you predict $A$, or it might be completely irrelevant, but it mathematically cannot increase your ignorance.

Correlations Never Shrink by Expanding the View: SSA can also be rearranged in terms of Quantum Mutual Information $I$, which measures the total correlation (both classical and quantum) between systems:

$$I(A:BC)\ge I(A:B)$$

where $I(A:B)$ is defined as $S_A+S_B-S_{AB}$. If $I(A:B) = 0$, it implies the two states share absolutely zero information, have completely independent probabilities, and form a strict, uncorrelated product state ($\rho_{AB} = \rho_A \otimes \rho_B$). Mutual Information answers a single physical question: How much does measuring subsystem $A$ reduce your ignorance about subsystem $B$?It quantifies the intersection, or the "shared" information, between two systems.$$I(A:B) = S(A) + S(B) - S(AB)$$You can read this equation conceptually as:(Individual Ignorance of A) + (Individual Ignorance of B) - (Total Joint Ignorance)If systems $A$ and $B$ are completely independent, knowing the state of $A$ tells you nothing about $B$. Your joint ignorance is just the sum of your individual ignorances ($S(AB) = S(A) + S(B)$). The subtraction perfectly cancels everything out, yielding $I(A:B) = 0$.

Getting back to the SSA, the total correlation that system $A$ shares with a combined system $BC$ must be at least as large as the correlation it shares with $B$ alone. If $A$ happens to be entangled with $C$, looking at the combined composite system $BC$ captures that extra link. You cannot lose correlations by expanding your observational boundary.

Physically, SSA is the strict mathematical engine that enforces the monogamy of quantum entanglement. This is exactly what creates the AMPS Firewall paradox we analyzed earlier.

Imagine a tripartite system where subsystem $B$ acts as a central node.

  • If $A$ and $B$ are in a maximally entangled pure state, they contain no joint uncertainty: $S(AB)=0$.

  • Because they are in a perfect pure state, their correlation budget is completely exhausted. System $B$ cannot share any entanglement or classical correlation with system $C$.

  • If $B$ is completely uncorrelated with $C$, then evaluating $BC$ is just adding their independent entropies: $S(BC)=S(B)+S(C)$.

Black Hole Thermodynamic Paradox

Since nothing can escape out of a black hole, if we look at it as a thermodynamic object, it means a black hole does not emit radiation. Therefore, as per Stefan-Boltzmann's Law for the power emitted by a blackbody at temperature $T$ and area $A$:

$$P = \sigma A T^4$$

Since $P = 0$, $T$ must be $0$. Thus, if some hot gas with heat energy $Q$ falls into the hole, the universe outside the hole lost entropy worth $Q/T$, where $T$ was the temperature of the gas. However, the black hole will have gained $Q/T_{\text{hole}}$ entropy. Since the temperature of the black hole is $0$, it means it gained an infinite amount of entropy which is obviously absurd. This is the classical black hole thermodynamic paradox. Heat will flow into the black hole, but never out. It can never reach equilibrium with a non-zero temperature bath.

Another way of looking at this is, $T = dQ_{\text{rev}} / dS$, specifically governs reversible heat transfer ($dQ_{\text{rev}}$). In classical general relativity, dropping matter or energy into a black hole is an irreversible process. Once energy crosses the event horizon, it cannot be extracted as heat. Because there is no reversible path to extract energy thermally from a classical black hole, the process $dQ_{\text{rev}}$ cannot be defined for the emission cycle. Therefore, one cannot use this identity to assign a temperature to the horizon.

The contradiction is resolved by proving that the black hole is not actually at $T=0$ by shifting from classical general relativity to quantum field theory in curved spacetime, where the entropy is related to the area of the black hole and its temperature is inversely proportional to its mass.

Rough Sketch of the Paradox

When a volume of gas with energy $Q$ and temperature $T_{\text{gas}}$ falls across the event horizon, we track the entropy change of the two systems independently. Note that to calculate the energy of the gas, we use the relativistic equation $E^2 = m^2 c^4 + p^2 c^2$.

The universe loses the gas. The entropy of the observable universe strictly decreases.

$$\Delta S_{\text{out}} = -\frac{Q}{T_{\text{gas}}}$$

The black hole absorbs the energy $Q$. By the mass-energy equivalence $dQ = c^2 dM$, the mass of the black hole increases. This increase in mass directly expands the surface area of the event horizon, thereby increasing the black hole's entropy.

Starting with the Schwarzschild radius $R_s = \frac{2GM}{c^2}$, the surface area $A$ of the spherical event horizon is:

$$A = 4\pi R_s^2 = \frac{16\pi G^2 M^2}{c^4}$$

Substitute this area into the Bekenstein-Hawking entropy formula:

$$S = \frac{k_B c^3}{4G\hbar} A = \frac{k_B c^3}{4G\hbar} \left( \frac{16\pi G^2 M^2}{c^4} \right) = \frac{4\pi k_B G M^2}{\hbar c}$$

Using the definition of temperature from the First Law of Thermodynamics, we divide the change in heat by the change in entropy:

$$T = \frac{dQ}{dS}$$

Substitute the expressions from Step 1 and Step 2:

$$T = \frac{c^2 dM}{\left( \frac{8\pi k_B G M}{\hbar c} \right) dM}$$

The $dM$ terms cancel out entirely, leaving:

$$T = \frac{\hbar c^3}{8\pi G M k_B}$$

This is the Hawking temperature $T_H$. Because the black hole acts as a thermal system at temperature $T_H$, its entropy increase is governed by:

$$\Delta S_{\text{BH}} = \frac{Q}{T_H}$$

We calculate the total change in entropy of the entire system:

$$\Delta S_{\text{total}} = \Delta S_{\text{BH}} + \Delta S_{\text{out}}$$

$$\Delta S_{\text{total}} = \frac{Q}{T_H} - \frac{Q}{T_{\text{gas}}}$$

For the gas to fall into the black hole and not be instantly vaporized and pushed outward by Hawking radiation, the temperature of the gas must be greater than the temperature of the black hole ($T_{\text{gas}} > T_H$).

Because $T_H$ is infinitesimally small compared to $T_{\text{gas}}$, the fraction $Q/T_H$ is massively larger than the fraction $Q/T_{\text{gas}}$.

The universe loses a small amount of entropy but the black hole gains an absolutely gargantuan (but strictly finite) amount of entropy. Therefore, $\Delta S_{\text{total}} > 0$.

The information and thermal entropy of the gas are not deleted; they are converted into the geometric entropy of the expanding event horizon. The total entropy of the universe increases.

The black hole gains a massive entropy but this manifests only as a small change in its area. The Bekenstein-Hawking entropy formula is defined as:

$$S_{\text{BH}} = \frac{k_B c^3}{4G\hbar} A$$

The constant fraction $\frac{c^3}{G\hbar}$ is the inverse of the Planck length squared ($l_P^2$). This allows the equation to be written geometrically as a ratio of areas:

$$S_{\text{BH}} = \frac{k_B}{4} \left( \frac{A}{l_P^2} \right)$$

The term $l_P^2$ represents the Planck area, which is approximately $2.6 \times 10^{-70} \text{ m}^2$. Because this exceptionally small number is in the denominator, the proportionality constant relating macroscopic square meters to entropy is immensely large (on the order of $10^{69} \text{ m}^{-2}$).

Looking back at the Stefan-Boltzmann law of radiation again from the non-classical point of view, because the black hole possesses a finite, non-zero temperature $T_H$, the Stefan-Boltzmann law dictates that it must radiate energy. This thermal blackbody emission is precisely what is known as Hawking radiation. The radiation originates strictly from just outside the event horizon, not from within the black hole.

The Quantum Information Paradox

Hawking’s thermodynamic solution created a fatal contradiction for quantum mechanics. Unitarity dictates that a pure quantum state ($S=0$) must evolve into another pure state. However, Hawking's calculations showed the emitted radiation is a perfectly thermal, mixed state ($S>0$), even if the blackhole initially formed from a pure state. If a black hole completely evaporates into a mixed thermal cloud, the initial quantum information (the specific wavefunctions of the matter that formed it) is permanently erased. Various mechanisms such as the Page Curve, AMPSS Postulates (the Firewall) and ENtanglement Islands were created to resolve this.

Aside: Note that a gas at some termperature in falling into a blackhole can be a mixed state due to the probability distribution of energy states of the gas, but the information paradox holds even when we theoretically assume a balckhole is formed from a pure state.

The Paradox in Terms of the AMPSS Postulates

The AMPSS (Almheiri, Marolf, Polchinski, Stanford, and Sully) argument streamlines the black hole information paradox into four postulates that appear reasonable based on our current understanding of black holes and quantum mechanics. However, AMPSS concluded that these four postulates cannot all be mutually consistent.

  • Unitarity: The formation and evaporation of a black hole is a unitary quantum mechanical process.

  • Local Effective Field Theory: Physics outside the horizon of a black hole is well-described by an effective local quantum field theory.

  • Quantum Black Holes: Black holes themselves are quantum mechanical systems possessing a discrete spectrum of states.

  • No Drama: For a sufficiently large black hole where local curvature at the horizon is very small, nothing special happens to an observer falling across the horizon into the black hole.

The Mathematical Setup

Suppose matter in a pure state collapses into a black hole of mass $M_0$. As it radiates, we collect all emitted radiation until the black hole's mass is substantially less than $M_0/2$, denoting the state of this early collected radiation as $\rho_R$.

We evaluate two specific modes of radiation:

  • $\rho_B$: A specific mode of radiation (a wave-packet) just outside the horizon that is leaving the black hole.

  • $\rho_A$: The partner mode just inside the horizon.

Based on the postulates and quantum field theory, we can establish four mathematical relationships:

  1. From Postulates 2 and 4: The joint state of $A$ and $B$ is entangled and pure.

    $$(i) \quad S(A)=S(B)\neq0, \quad S(AB)=0$$

  2. From Subadditivity of Entanglement Entropy: Because $\rho_{AB}$is pure, it forms a tensor product with the rest of the system such that$\rho_{ABR}=\rho_{AB}\otimes\rho_R$.

    $$(ii) \quad S(ABR)=S(R)$$

    1. From Postulate 1 (Unitarity): Once the black hole has evaporated past half its initial mass, any newly emitted quantum of radiation should purify the earlier emitted radiation.

    $$(iii) \quad S(BR)<S(R)$$

  3. From Strong Subadditivity: Finally, we apply the strong subadditivity of entropy among the $A$, $B$, and $R$ subsystems:

    $$(iv) \quad S(AB) + S(BR) \ge S(ABR) + S(B)$$

    (Note: The condition $S(BR) < S(R)$ from postulate 1 holds more precisely past the Page time, at which point the black hole's horizon area reaches half its initial value, meaning the radiation $R$ constitutes the majority of the system).

The Mathematical Contradiction

Putting it all together, we can construct the following logical chain:

$$S(R) + S(B) = S(ABR) + S(B) \quad \text{using (ii)}$$

$$S(ABR) + S(B) \le S(AB) + S(BR) \quad \text{using (iv)}$$

$$S(AB) + S(BR) < S(AB) + S(R) \quad \text{using (iii)}$$

$$S(AB) + S(R) = S(R) \quad \text{using (i)}$$

Tracing the inequality from the beginning to the end of the chain yields:

$$S(R) + S(B) < S(R)$$

For this to be true, $S(B)$ must be strictly less than zero. However, since $B$ is a quantum subsystem, its entropy must be non-negative, and by postulate (i), $S(B) \neq 0$. We have therefore arrived at a strict mathematical contradiction.

Consequences & Resolutions

AMPSS' conclusion was that at least one of their four postulates has to be modified. How palatable the ensuing consequences are is up to you to reason through.

  1. Drop Unitarity: If we drop unitarity entirely, we concede that black holes fundamentally destroy quantum information.

  2. Modify Local Effective Field Theory: One way to modify local effective field theory is to delete the word "local" and allow for small amounts of nonlocality. Holographic resolutions of the black hole information problem (and AdS/CFT itself) are inherently nonlocal in the sense that degrees of freedom are replicated in both the bulk space-time and its boundary.

  3. Remnants or Non-Quantum Black Holes: One way to evade the AMPSS argument is if black holes never finish evaporating, instead leaving behind a small and extremely entropic remnant. Alternatively, perhaps black holes are simply not described by quantum mechanics at all.

  4. The Firewall (Drop "No Drama"): If $A$ and $B$ are not in a pure entangled state such that $S(AB) \neq 0$, then it is possible to evade the contradiction. However, such states possess massive local energy densities. In this setting, it would be as if there was a "firewall" waiting just behind the horizon that an infalling observer would hit as they entered the black hole. As AMPSS pointed out, the result is considerable drama for the observer.

What does AdS/CFT say about the Paradox

Within the holographic AdS/CFT framework, the paradox is "resolved" by sacrificing strict classical locality. The core mathematical mechanism is that the Hilbert space of the black hole interior (subsystem $A$) and the Hilbert space of the distant Hawking radiation (subsystem $R$) are not entirely distinct tensor factors.

Because the boundary of an AdS space-time completely encodes its bulk geometry, the distant radiation actually contains the degrees of freedom of the interior. By dynamically encoding $A$ within $R$, the contradiction regarding strong subadditivity is bypassed, and unitary evolution is preserved. However, this is considered a "resolution" only in negatively curved anti-de Sitter space; proving it works in our flat space-time remains unsolved.

The Problem with Dropping Locality

Locality is the foundational mathematical pillar of both General Relativity and standard Quantum Field Theory. We are extremely hesitant to abandon it for two reasons:

  • Causality: In a local effective field theory, operators at space-like separations must commute: $[O(x), O(y)] = 0$. If an event at $x$ can dynamically affect $y$ across space-like intervals, it opens the door to faster-than-light information transfer and causal loops.

  • Cluster Decomposition: The S-matrix in QFT relies on the principle that distant experiments must yield independent results.

Sacrificing locality threatens the mathematical scaffolding of the entire Standard Model. Accepting nonlocality requires proving that it is confined strictly to the extreme gravitational gradients of quantum gravity, preventing it from bleeding out and breaking classical relativistic observables.

Current Status of the Information Paradox

The theoretical consensus leans heavily toward modifying locality to save quantum unitarity, though the exact physical geometry remains under active research. The leading frameworks include:

  • Entanglement Islands: Post-2019 calculations using the replica trick demonstrate that the interior of the black hole (the "island") mathematically belongs to the entanglement wedge of the distant Hawking radiation. Once the black hole passes the Page time, measuring the distant radiation implicitly measures the interior.

  • ER=EPR Conjecture: This proposes that quantum entanglement (Einstein-Podolsky-Rosen) and spacetime wormholes (Einstein-Rosen bridges) are exactly the same phenomenon. The entangled interior mode $A$ and the early radiation $R$ are connected via microscopic geometric wormholes, bypassing the local space-time barrier of the event horizon entirely.

Status of Proof for De Sitter Space

The entanglement islands proposal was originally developed and rigorously proven within Anti-de Sitter (AdS) space, but it has since become a major focal point for research in de Sitter (dS) space and flat spacetime.

The semi-classical calculations demonstrate that entanglement islands do appear in de Sitter space, but their physical implications differ from those in AdS black holes:

  • Cosmological Horizons: In dS space, an observer is surrounded by a cosmological horizon (the boundary beyond which light can never reach us due to cosmic expansion). Like a black hole event horizon, this cosmological horizon emits thermal Gibbons-Hawking radiation.

  • The Island Location: Calculations show that if you collect the Gibbons-Hawking radiation for a sufficiently long time, an entanglement island forms. However, this island is located outside the cosmological horizon, in the region of space causally disconnected from the observer.

  • Black Holes in dS: If you place a black hole inside a de Sitter universe, the math shows that the island formula still resolves the black hole information paradox, yielding a unitary Page curve just as it did in AdS.

While the semi-classical replica trick successfully generates islands in dS space, these results are widely considered less rigorous than the AdS derivations.

The primary issue is the lack of a strict dS/CFT correspondence. In AdS, the conformal boundary is time-like, meaning it sits at a fixed spatial infinity where we can safely define a non-gravitating quantum system. In de Sitter space, the conformal boundary is space-like (it exists in the infinite future, $t \to \infty$).

Because there is no stable, time-like boundary in dS space where we can stand to collect and measure the Hawking radiation, defining the exact Hilbert space of the "exterior radiation" requires ad-hoc mathematical boundaries (often artificially coupling the dS space to a flat auxiliary bath).

Petz Maps

For a noisy channel $\mathcal{N}$, if $D(\rho\vert{}\vert{}\sigma)=D(\mathcal{N}(\rho)\vert{}\vert{}\mathcal{N}(\sigma))$, where $D(\rho\vert{}\vert{}\sigma)$ is the relative entropy given by $\text{Tr}(\rho\log\rho-\rho\log\sigma)$, it means the distinguishability of the two states does not reduce when passed through the channel. Under this condition, the noisy channel is perfectly reversible on these states.

If there is a $\sigma\in\mathcal{Q}$ such that $\text{supp}(\rho)\subseteq\text{supp}(\sigma)$ for all $\rho\in\mathcal{Q}$, a channel that undoes the action of $\mathcal{N}$ is the Petz map, defined as:

$$\mathcal{P}_{\sigma,\mathcal{N}}(\rho)=\sigma^{1/2}\mathcal{N}^\dagger\left(\mathcal{N}(\sigma)^{-1/2}\rho\mathcal{N}(\sigma)^{-1/2}\right)\sigma^{1/2}$$

$\sigma$ is a specifically chosen baseline state such that $\text{supp}(\rho)\subseteq\text{supp}(\sigma)$ for all $\rho$ in the space. The "support" of a density matrix is the subspace spanned by eigenvectors corresponding to non-zero eigenvalues. This condition mandates that the kernel of $\sigma$ is a subset of the kernel of $\rho$, and therefore $\sigma$ possesses a non-zero probability anywhere that any other state $\rho$ in the set might possess a non-zero probability. This is essential for the notion of relative entropy to not blow to infinity. If $\sigma$ has a zero eigenvalue for some eigenstate where $\rho$ has a non-zero eigenvalue, the formula for relative entropy will blow up due to the factor of $\rho$ being non-zero when $\log\sigma$ hits negative infinity.

Aside: The log of a matrix is calculated by taking the log of the eigenvalues and reconstructing the matrix using the spectral theorem.

Here is the rigorous construction of the Petz map for the 3-qubit repetition code for the bit flip error. This derivation demonstrates exactly how the Petz map bypasses the need for explicit classical syndrome measurements.

Let the code space $\mathcal{C}$ be the subspace spanned by the logical basis ${\vert{}000\rangle,\vert{}111\rangle}$. Let $P_0=\vert{}000\rangle\langle000\vert{}+\vert{}111\rangle\langle111\vert{}$ be the projector onto this space.

  • The Target State ($\rho$): The specific encoded quantum information we want to protect, $\rho=\vert{}\psi\rangle\langle\psi\vert{}$, where $\vert{}\psi\rangle=\alpha\vert{}000\rangle+\beta\vert{}111\rangle$.

  • The Reference State ($\sigma$): To satisfy the support condition ($\text{supp}(\rho)\subseteq\text{supp}(\sigma)$), we choose the maximally mixed state within the code space:

  • $$\sigma=\frac{1}{2}P_0=\frac{1}{2}(\vert{}000\rangle\langle000\vert{}+\vert{}111\rangle\langle111\vert{})$$

  • The Noise Channel ($\mathcal{N}$): Let the environment apply a bit-flip (Pauli $X_1$) to the first qubit with probability $p$. The forward channel is defined by Kraus operators $E_0=\sqrt{1-p}I$ and $E_1=\sqrt{p}X_1$:

    $$\mathcal{N}(\rho)=(1-p)\rho+pX_1\rho X_1$$

To build the Petz map, we first evaluate how the noise channel distorts the reference state $\sigma$.

$$\mathcal{N}(\sigma)=\frac{1-p}{2}P_0+\frac{p}{2}X_1P_0X_1$$

Let $P_1=X_1P_0X_1=\vert{}100\rangle\langle100\vert{}+\vert{}011\rangle\langle011\vert{}$ be the projector onto the error subspace. Because a bit flip strictly moves the logical states out of the code space, $P_0$ and $P_1$ are mutually orthogonal ($P_0P_1=0$).

$$\mathcal{N}(\sigma)=\frac{1-p}{2}P_0+\frac{p}{2}P_1$$

We require the inverse square root of this operator on its support (the pseudoinverse). Because $P_0$ and $P_1$ are orthogonal projectors, we simply invert the square roots of their coefficients:

$$\mathcal{N}(\sigma)^{-1/2}=\sqrt{\frac{2}{1-p}}P_0+\sqrt{\frac{2}{p}}P_1$$

We apply this normalization to the degraded target state $\mathcal{N}(\rho)$. This evaluates the inner bracket of the Petz map:

$$Y=\mathcal{N}(\sigma)^{-1/2}\left[\mathcal{N}(\rho)\right]\mathcal{N}(\sigma)^{-1/2}$$

Substitute the noisy state $\mathcal{N}(\rho)=(1-p)\rho+pX_1\rho X_1$. Because $\rho$ is strictly supported on $P_0$ and $X_1\rho X_1$ is strictly supported on $P_1$, all cross-multiplication terms between the orthogonal projectors evaluate to zero:

$$Y=\left(\sqrt{\frac{2}{1-p}}P_0\right)(1-p)\rho\left(\sqrt{\frac{2}{1-p}}P_0\right)+\left(\sqrt{\frac{2}{p}}P_1\right)pX_1\rho X_1\left(\sqrt{\frac{2}{p}}P_1\right)$$

The noise probabilities $(1-p)$ and $p$ perfectly cancel out with the normalization scalars:

$$Y=2\rho+2X_1\rho X_1$$

Next, we apply the adjoint map $\mathcal{N}^\dagger$ to the operator $Y$. The adjoint of our channel is $\mathcal{N}^\dagger(Y)=E_0^\dagger YE_0+E_1^\dagger YE_1=(1-p)Y+pX_1YX_1$.

Substitute $Y$:

$$\mathcal{N}^\dagger(Y)=(1-p)(2\rho+2X_1\rho X_1)+pX_1(2\rho+2X_1\rho X_1)X_1$$

Because Pauli operators square to the identity ($X_1^2=I$), distributing the $X_1$ on the right side yields $2X_1\rho X_1+2\rho$.

$$\mathcal{N}^\dagger(Y)=(1-p)(2\rho+2X_1\rho X_1)+p(2X_1\rho X_1+2\rho)$$

$$\mathcal{N}^\dagger(Y)=(1-p+p)(2\rho+2X_1\rho X_1)=2\rho+2X_1\rho X_1$$

The Heisenberg evolution of the adjoint channel perfectly maps the operator back to itself.

The final step of the Petz map rotates the state back into the original Hilbert subspace by sandwiching it with $\sigma^{1/2}$.

$$\sigma^{1/2}=\frac{1}{\sqrt{2}}P_0$$

$$\mathcal{P}_{\sigma,\mathcal{N}}(\mathcal{N}(\rho))=\left(\frac{1}{\sqrt{2}}P_0\right)\left[2\rho+2X_1\rho X_1\right]\left(\frac{1}{\sqrt{2}}P_0\right)$$

$$\mathcal{P}_{\sigma,\mathcal{N}}(\mathcal{N}(\rho))=\frac{1}{2}P_0(2\rho)P_0+\frac{1}{2}P_0(2X_1\rho X_1)P_0$$

Since $P_0X_1\rho X_1P_0=0$, the error term is mathematically annihilated. Furthermore, since $\rho$ lives entirely within the $P_0$ subspace, $P_0\rho P_0=\rho$.

$$\mathcal{P}_{\sigma,\mathcal{N}}(\mathcal{N}(\rho))=\rho$$

The initial state is perfectly recovered. The matrix inversion mapped the error probability distribution, the adjoint channel pulled the physical observables backward through the noise, and the final $\sigma^{1/2}$ weighting executed the exact equivalent of conditional error correction without ever needing a measurement circuit.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment