Un’appendice culturale che collega la probabilità all’informazione. L’entropia misura il “disordine” o l’incertezza di una variabile aleatoria: la stessa formula H=plogpH=-\sum p\log p descrive il numero medio minimo di bit necessari per identificare un esito e ha la stessa struttura dell’entropia termodinamica di Boltzmann. Claude Shannon la introdusse negli anni ‘40 come misura dell’informazione media trasmessa da una sorgente probabilistica.

A cultural appendix that connects probability to information. Entropy measures the “disorder” or the uncertainty of a random variable: the same formula H=plogpH=-\sum p\log p describes the minimum average number of bits needed to identify an outcome and has the same structure as Boltzmann’s thermodynamic entropy. Claude Shannon introduced it in the 1940s as a measure of the average information transmitted by a probabilistic source.

L’entropia è una grandezza che misura il “disordine”. In termodinamica la si introduce come “tasso di cambio” tra energia e disordine: ΔE=TΔS\Delta E = T\Delta S (Borghi, Entropioni, 2019). Claude Shannon negli anni ‘40 scoprì che la stessa formula descrive l’informazione media trasmessa da una sorgente probabilistica (Shannon, 1948).

Definizione — Entropia di Shannon

Sia XX una variabile aleatoria discreta con valori x1,,xnx_1,\ldots,x_n e probabilità p1,,pnp_1,\ldots,p_n (con pi=1\sum p_i=1). L’entropia di XX è H(X)=i=1npilog2pi,H(X) = -\sum_{i=1}^{n} p_i \log_2 p_i, misurata in bit. Convenzione: 0log0=00\cdot\log 0 = 0 (limite di plogpp\log p per p0+p\to 0^+). Cambiando la base del logaritmo cambia l’unità: base eenat, base 1010hartley.

Collegamenti

Argomenti: Distribuzioni probabilita
Concetti: Entropia · Variabile aleatoria
Competenze: Modellizzare
Persone: Shannon

Entropy is a quantity that measures “disorder”. In thermodynamics it is introduced as the “exchange rate” between energy and disorder: ΔE=TΔS\Delta E = T\Delta S (Borghi, Entropioni, 2019). Claude Shannon in the 1940s discovered that the same formula describes the average information transmitted by a probabilistic source (Shannon, 1948).

Definition — Shannon entropy

Let XX be a discrete random variable with values x1,,xnx_1,\ldots,x_n and probabilities p1,,pnp_1,\ldots,p_n (with pi=1\sum p_i=1). The entropy of XX is H(X)=i=1npilog2pi,H(X) = -\sum_{i=1}^{n} p_i \log_2 p_i, measured in bits. Convention: 0log0=00\cdot\log 0 = 0 (limit of plogpp\log p as p0+p\to 0^+). Changing the base of the logarithm changes the unit: base ee gives nats, base 1010 gives hartleys.

Topics: Distribuzioni probabilita
Concepts: Entropia · Variabile aleatoria
Skills: Modellizzare
People: Shannon

La definizione di entropia diventa concreta se la si legge come una domanda: quante domande “sì / no” servono in media per scoprire il valore di XX?

Osservazione — Significato operativo

H(X)H(X) è il numero medio minimo di domande sì / no (o di bit, in una codifica binaria ottimale) necessari per identificare il valore di XX. Distribuzione uniforme su 2k2^k valori: H=kH = k bit (servono esattamente kk domande). Distribuzione deterministica: H=0H = 0 (non c’è nulla da scoprire).

Questa lettura spiega perché l’entropia si misura in bit: un bit è esattamente la risposta a una domanda binaria ben posta. Più una distribuzione è “prevedibile”, meno domande servono, e più bassa è la sua entropia.

Collegamenti

Argomenti: Distribuzioni probabilita
Concetti: Entropia
Competenze: Interpretare grafico

The definition of entropy becomes concrete if you read it as a question: how many “yes / no” questions are needed on average to discover the value of XX?

Remark — Operational meaning

H(X)H(X) is the minimum average number of yes / no questions (or of bits, in an optimal binary encoding) needed to identify the value of XX. Uniform distribution over 2k2^k values: H=kH = k bits (exactly kk questions are needed). Deterministic distribution: H=0H = 0 (there is nothing to discover).

This reading explains why entropy is measured in bits: a bit is exactly the answer to a well-posed binary question. The more “predictable” a distribution, the fewer questions are needed, and the lower its entropy.

Topics: Distribuzioni probabilita
Concepts: Entropia
Skills: Interpretare grafico

Il caso più semplice — una moneta eventualmente truccata — mostra come l’entropia sia una funzione della probabilità pp, con un massimo ben preciso.

Esempio — Lancio di una moneta

X{T,C}X\in\{T,C\} con P(T)=pP(T)=p, P(C)=1pP(C)=1-p. Allora H(X)=plog2p(1p)log2(1p).H(X) = -p\log_2 p - (1-p)\log_2(1-p). Studio di funzione: massimo in p=1/2p=1/2 con H=1H=1 bit (totale incertezza); minimo in p=0p=0 o p=1p=1 con H=0H=0 (nessuna sorpresa).

Entropia della moneta H(p)=plog2p(1p)log2(1p)H(p)=-p\log_2 p-(1-p)\log_2(1-p): massimo Hmax=1H_{\max}=1 bit in p=1/2p=1/2, nulla agli estremi p=0p=0 e p=1p=1.

Collegamenti

Argomenti: Distribuzioni probabilita
Concetti: Entropia
Competenze: Interpretare grafico · Studiare funzione

The simplest case — a possibly biased coin — shows how entropy is a function of the probability pp, with a well-defined maximum.

Example — Tossing a coin

X{T,C}X\in\{T,C\} with P(T)=pP(T)=p, P(C)=1pP(C)=1-p. Then H(X)=plog2p(1p)log2(1p).H(X) = -p\log_2 p - (1-p)\log_2(1-p). Function study: maximum at p=1/2p=1/2 with H=1H=1 bit (total uncertainty); minimum at p=0p=0 or p=1p=1 with H=0H=0 (no surprise).

Entropy of the coin H(p)=plog2p(1p)log2(1p)H(p)=-p\log_2 p-(1-p)\log_2(1-p): maximum Hmax=1H_{\max}=1 bit at p=1/2p=1/2, zero at the extremes p=0p=0 and p=1p=1.

Topics: Distribuzioni probabilita
Concepts: Entropia
Skills: Interpretare grafico · Studiare funzione

Con più di due esiti l’entropia cresce, ma dipende da quanto la distribuzione è “sbilanciata”: un dado equo è più imprevedibile di uno truccato.

Esempio — Dado a sei facce

XX uniforme su {1,,6}\{1,\ldots,6\}: H=616log216=log262,585 bit.H = -6\cdot\frac{1}{6}\log_2\frac{1}{6} = \log_2 6 \approx 2{,}585 \text{ bit}. Servono in media circa 2,62{,}6 domande sì / no ben scelte per indovinare l’esito.

Dado truccato con P(6)=1/2P(6)=1/2 e P(1)==P(5)=1/10P(1)=\cdots=P(5)=1/10: H=12log2125110log2110=12+12log2102,16 bit.H = -\tfrac12\log_2\tfrac12 - 5\cdot\tfrac{1}{10}\log_2\tfrac{1}{10} = \tfrac12 + \tfrac12\log_2 10 \approx 2{,}16 \text{ bit}. Sapere che il 66 è probabile riduce l’incertezza: meno bit per descrivere il risultato medio.

Il confronto illustra un principio generale: a parità di numero di esiti, la distribuzione uniforme è quella di entropia massima; ogni squilibrio nelle probabilità abbassa l’entropia.

Collegamenti

Argomenti: Distribuzioni probabilita
Concetti: Entropia
Competenze: Calcolare · Interpretare grafico

With more than two outcomes the entropy grows, but it depends on how “unbalanced” the distribution is: a fair die is more unpredictable than a biased one.

Example — Six-sided die

XX uniform over {1,,6}\{1,\ldots,6\}: H=616log216=log262,585 bit.H = -6\cdot\frac{1}{6}\log_2\frac{1}{6} = \log_2 6 \approx 2{,}585 \text{ bit}. On average about 2,62{,}6 well-chosen yes / no questions are needed to guess the outcome.

Biased die with P(6)=1/2P(6)=1/2 and P(1)==P(5)=1/10P(1)=\cdots=P(5)=1/10: H=12log2125110log2110=12+12log2102,16 bit.H = -\tfrac12\log_2\tfrac12 - 5\cdot\tfrac{1}{10}\log_2\tfrac{1}{10} = \tfrac12 + \tfrac12\log_2 10 \approx 2{,}16 \text{ bit}. Knowing that the 66 is likely reduces the uncertainty: fewer bits to describe the average result.

The comparison illustrates a general principle: for a given number of outcomes, the uniform distribution is the one of maximum entropy; any imbalance in the probabilities lowers the entropy.

Topics: Distribuzioni probabilita
Concepts: Entropia
Skills: Calcolare · Interpretare grafico

Tre proprietà racchiudono il comportamento dell’entropia: è sempre non negativa, ha un massimo ben preciso, e si somma per variabili indipendenti.

Proprietà — Proprietà fondamentali

  • H(X)0H(X)\ge 0, con =0=0 se e solo se XX è costante;
  • H(X)log2nH(X)\le \log_2 n, con == se e solo se XX è uniforme (l’uniforme massimizza l’entropia, intuitivamente: la massima ignoranza);
  • additività per indipendenza: H(X,Y)=H(X)+H(Y)H(X,Y) = H(X) + H(Y) se X,YX,Y sono indipendenti.

La seconda proprietà è la versione rigorosa dell’osservazione sul dado: fra tutte le distribuzioni su nn esiti, quella uniforme è la più imprevedibile e la sua entropia log2n\log_2 n è il massimo possibile.

Collegamenti

Argomenti: Distribuzioni probabilita
Concetti: Entropia
Competenze: Dimostrare

Three properties capture the behaviour of entropy: it is always non-negative, it has a well-defined maximum, and it adds up for independent variables.

Property — Fundamental properties

  • H(X)0H(X)\ge 0, with =0=0 if and only if XX is constant;
  • H(X)log2nH(X)\le \log_2 n, with == if and only if XX is uniform (the uniform maximises entropy, intuitively: maximum ignorance);
  • additivity under independence: H(X,Y)=H(X)+H(Y)H(X,Y) = H(X) + H(Y) if X,YX,Y are independent.

The second property is the rigorous version of the observation about the die: among all distributions over nn outcomes, the uniform one is the most unpredictable and its entropy log2n\log_2 n is the maximum possible.

Topics: Distribuzioni probabilita
Concepts: Entropia
Skills: Dimostrare

Il nome “entropia” scelto da Shannon non è un caso: la sua formula ha la stessa struttura di quella della termodinamica, e il ponte tra i due mondi è uno dei fili conduttori più affascinanti della matematica applicata.

Osservazione — Dal disordine termodinamico al disordine informatico

La formula di Shannon ha la stessa struttura della formula di Boltzmann S=kBlnWS = k_B \ln W (con WW = numero di microstati equiprobabili). Per questo motivo Shannon, su suggerimento di von Neumann, la chiamò “entropia” con la celebre giustificazione: “vincerai sempre i dibattiti perché nessuno sa davvero cosa sia l’entropia” (Shannon, 1948). È il filo conduttore di Borghi (Entropioni): la stessa grandezza descrive trasferimenti reversibili di energia, ottimalità dei codici di compressione, e l’incertezza di un esperimento aleatorio.

Collegamenti

Argomenti: Distribuzioni probabilita
Concetti: Entropia
Persone: Boltzmann · Shannon · Von Neumann

The name “entropy” chosen by Shannon is no coincidence: his formula has the same structure as the one in thermodynamics, and the bridge between the two worlds is one of the most fascinating threads of applied mathematics.

Remark — From thermodynamic disorder to information disorder

Shannon’s formula has the same structure as Boltzmann’s formula S=kBlnWS = k_B \ln W (with WW = number of equiprobable microstates). For this reason Shannon, on von Neumann’s suggestion, called it “entropy” with the celebrated justification: “you will always win debates because nobody really knows what entropy is” (Shannon, 1948). It is the guiding thread of Borghi (Entropioni): the same quantity describes reversible transfers of energy, the optimality of compression codes, and the uncertainty of a random experiment.

Topics: Distribuzioni probabilita
Concepts: Entropia
People: Boltzmann · Shannon · Von Neumann