Why this matters

If you operate a streaming service, an OTT platform, an e-learning library, or a video-conferencing product, your CDN bill scales linearly with bitrate and your viewer experience scales non-linearly with quality. Every percentage point a new codec saves at the same quality is real money. But you cannot plan around vendor marketing: a slide claiming "70 percent better than AV1" needs to be read against the theoretical maximum any codec can ever achieve, otherwise you size your encoding farm for a number that will never arrive. This article gives you the tools to read those claims — what Shannon's limit actually is, what it does and doesn't bound, where conventional and neural codecs sit on the curve in 2026, and how to estimate, for your own content and your own quality target, how much room is left.

What Shannon actually said in 1948

Claude Shannon's 1948 paper A Mathematical Theory of Communication did two related things. First, he defined a number called entropy, written H, that measures the average information content of a source — the average number of bits per symbol that the source produces, given how its symbols are distributed.1 The phrase "average information content" is doing real work here. If a source always emits the same symbol, its entropy is zero — the sequence is fully predictable, so describing it costs zero bits per symbol on average. If a source emits one of 256 symbols with equal probability, its entropy is 8 bits — every symbol is maximally surprising.

Second, Shannon proved the source coding theorem: it is impossible to compress a source losslessly below its entropy without losing information, and any compression scheme that uses fewer bits per symbol than H must, on average, fail to reconstruct the source.2 You can get arbitrarily close to H — modern entropy coders like arithmetic coding and CABAC sit within a fraction of a percent of the limit — but you cannot go beneath it. That is the lossless floor.

A useful analogy. A locked filing cabinet has a key, and the key has a minimum number of bits required to describe its shape. You can package the key in a smaller envelope by folding it cleverly, but the key itself is exactly N bits long. You cannot shrink the key — only the envelope.

For video, the lossless floor matters in archival and medical imaging, where formats like FFV1, ProRes 4444, and JPEG-XL operate, but it is not where consumer codecs live. Consumer codecs are lossy, which means they get to throw information away — and that opens a different door.

The lossy door: rate-distortion theory

In the same body of work, Shannon defined a second limit that does apply to consumer video. The rate-distortion function R(D) gives the smallest possible bitrate R at which a source can be encoded such that the average reconstruction error stays below a target distortion D.3 R(D) is a curve, not a single number. As you allow more distortion (D rises), the required bitrate (R) falls. As you demand near-perfect reconstruction (D approaches zero), R climbs toward the entropy H.

The shape of the R(D) curve depends on three things. First, what the source actually looks like — a low-motion talking-head shot has a different curve from a 4K sports broadcast, and a piece of computer-generated animation has yet another. Second, how you measure distortion — mean squared error gives one curve, structural similarity gives another, VMAF gives a third. Third, the statistical model of the source — a richer model, capturing temporal and spatial correlations, produces a tighter (lower) R(D) curve.

R(D) is the right object to think about when you compare codecs. Every codec, conventional or neural, can be plotted as a point or a curve in (R, D) space. The Shannon limit is the lower-left envelope: the curve below which no codec can go for that particular source and that particular distortion measure. Generational progress in codecs is the slow walk of the codec curve down toward the Shannon limit.

A line chart titled Codec generations walk toward the R(D) floor. The horizontal axis is bitrate at the same quality relative to H.264 on a log scale, marked 0.05, 0.1, 0.2, 0.5, 1 and 2; the vertical axis is quality, low to high. Six curves rise from left to right, each with its own colour and dash pattern: a thick solid Shannon limit R(D) on top, then DCVC-FM (2024), VVC (2020), AV1 (2018), HEVC (2013) and H.264 (2003). The empty area above the top curve is labelled No codec can reach this region. A dashed horizontal line labelled Same quality target crosses all six, and a dot marks each crossing: the Shannon limit at 0.09, DCVC-FM at 0.18, VVC at 0.25, AV1 at 0.35, HEVC at 0.50 and H.264 at 1.0, so each generation reaches the same quality further to the left. A note says the curves are schematic, that each codec sits at its measured BD-rate saving against H.264 at the same quality, that DCVC-FM is neural while the other four are conventional standards, and that the gap to the floor shrinks each generation, and so does the room left to gain Figure 1. Codec progress reads as a walk toward the rate-distortion floor. Each generation closes part of the remaining gap, but the easy gains are behind us.

Where conventional codecs sit on the curve in 2026

The most-cited generational milestones in commercial codecs measure relative bitrate savings using the Bjøntegaard delta rate, abbreviated BD-rate. BD-rate gives the average bitrate change between two codecs at the same objective quality, integrated across a range of quality points.4 The numbers below are BD-rate savings against H.264, measured on standard JVET test sets, mostly at HD and UHD resolutions.

HEVC (2013) buys roughly 50 percent over H.264 on standardised test content, with the win larger on UHD than HD because HEVC's larger coding tree units handle high-resolution flat areas more efficiently.5 AV1 (2018) lands at roughly 30 percent over HEVC on the same content, or about 65 percent over H.264.6 VVC (2020) measures about 40 to 50 percent over HEVC on UHD content, which compounds to roughly 75 percent over H.264.7 At 8K resolution, BBC R&D measurements published in late 2024 land VVC at 78 percent savings, AV1 at 63 percent, and HEVC at 53 percent — all relative to H.264.8

AV2, finalised in late 2025, adds another 30 to 40 percent over AV1 on Google's internal test sets.9 That compounds to roughly 77 percent over H.264 — almost identical to VVC. AOMedia and the JVET project (which produces VVC and the experimental ECM extensions) are now within a few percentage points of each other on most content.

A useful pattern emerges. H.264 to HEVC was a 50 percent halving. HEVC to AV1 was a 30 percent third. AV1 to AV2 is another 30 percent third. The absolute gain is shrinking each generation, but more importantly the gap that remains is shrinking faster — which means each future generation has less room to claim. The Shannon limit for natural video, on conventional distortion measures, is not infinitely far away.

Codec Year BD-rate vs H.264 Notes
H.264 / AVC 2003 0% (baseline) Workhorse of the internet
HEVC / H.265 2013 -50% Patent thicket slowed adoption
AV1 2018 -65% Royalty-free; 30% of Netflix viewing in 2025
VVC / H.266 2020 -75% Best conventional standard; weak market traction
AV2 2025 -77% ~30% over AV1; AOMedia finalised late 2025

The pattern is the diminishing-returns curve of any maturing technology. We are not at the floor — there is still room, especially on high-motion and high-resolution content — but the room is thinner each cycle.

The new dimension: perception

For a long time, the codec community measured progress along the (R, D) axes using objective distortion measures like PSNR or SSIM. The dirty secret of those measures is that they correlate poorly with what a human viewer actually sees. A picture can have low MSE and look terrible (waxy, over-smoothed faces) or high MSE and look great (realistic film grain, sharp textures). Netflix's VMAF was a major step forward because it correlates better with subjective scores, but it is still a per-frame regression onto a fixed observer model.

In 2018 and 2019, Yochai Blau and Tomer Michaeli published a series of papers that re-framed the problem in a way that has become central to the field.10 They showed that distortion and perceptual quality are mathematically at odds — at a fixed bitrate, you cannot simultaneously minimise the mean squared error and the divergence between the distribution of reconstructed images and the distribution of natural images. The two objectives trade off, and the trade-off has a precise rate-distortion-perception (RDP) function.11

The practical consequence: at a given bitrate, the picture can be optimised to look correct ("plausible image from the same distribution") at the cost of being slightly wrong ("each pixel differs from the original more than necessary"). Or it can be optimised to be pixel-accurate, at the cost of looking subtly degraded. The two cannot both be minimised. Modern neural codecs, which can be trained against any differentiable loss, have started to exploit this — they sit on a different part of the RDP surface than conventional codecs do.

A two-dimensional chart titled The rate-distortion-perception trade-off, subtitled that at one fixed bitrate a codec can be pixel-accurate or look natural, not both. The horizontal axis is per-pixel distortion D, low to high; the vertical axis is perceptual divergence P, low to high; a note says bitrate R is held fixed. One convex curve falls from the top left to the bottom right. The shaded area below and left of the curve is labelled Not reachable at this bitrate, and the open area above it is labelled Reachable, but worse on both axes. A blue dot high on the left is labelled Conventional target, HEVC, AV1, VVC, AV2; a purple dot low on the right is labelled Neural target, DCVC-FM, DCVC-RT. Two panels below carry the same point: conventional codecs optimise the per-pixel error, so the picture matches the source but can look waxy, while neural codecs match the distribution of natural video, so it looks right and each pixel differs more Figure 2. The Blau-Michaeli rate-distortion-perception surface. Conventional codecs optimise distortion; neural codecs can shift along the perception axis at the same bitrate.

Where neural codecs stand in 2026

A neural video codec replaces some or all of the hand-designed blocks in a conventional encoder — transform, prediction, entropy coder — with a learned model. The learned model is trained directly on the rate-distortion (or rate-distortion-perception) objective using stochastic gradient descent, which means it can exploit statistical regularities in natural video that hand-coded blocks miss.

Microsoft Research's DCVC family is the most-cited line of work. DCVC-DC (2023) was the first neural video codec to beat VVC's reference implementation VTM on standard JVET test sets under intra-period 32. DCVC-FM (2024), with feature modulation and multi-scale temporal context, beat VTM by 25.5 percent in BD-rate on the same test set under the more demanding intra-period -1 setting (a single intra frame at the start, then unlimited inter prediction).12 DCVC-RT (CVPR 2025) is the first real-time neural codec — it reaches roughly 21 percent BD-rate gain over VVC at framerates that match conventional encoders.13

A more recent benchmark, published at CVPR 2025, evaluated DCVC-FM, DCVC-DC, ECM (the JVET experimental codec that sits above VVC), and AVM (AOM's experimental codec that became AV2). Measured by VMAF on 4K content, DCVC-FM achieves bitrate savings of more than 8 percent over AVM, comparable to ECM, and over 37 percent over HEVC.14 Measured by perceptual quality more carefully — using LPIPS or FID against reference video distributions — neural codecs pull further ahead.

The catch in 2026 is the same catch neural codecs have had since they appeared: compute. A neural decoder requires a GPU or a substantial NPU, runs at orders-of-magnitude higher power than a hardware HEVC decoder, and is not yet present in any consumer device's silicon. DCVC-RT's "real-time" benchmarks run on an NVIDIA RTX 4090 — a 450-watt desktop card. That is fine for VOD pre-encoding (where you spend GPU once and recoup it on every play) and for cloud transcoding, but it is not yet a delivery codec.

How far is the Shannon limit, really?

A useful way to estimate is to compare the best current codec to a lower bound on R(D) computed from the source. For natural video at moderate compression rates, several independent estimates land in similar territory. Theoretical work on Gaussian models suggests that the remaining gap from VVC and DCVC-FM to the unconstrained R(D) bound is in the single-digit percentage range at typical streaming bitrates, and somewhat larger at very-low-bitrate operating points.15 In other words, for the workhorse 1080p and 4K rates where most of the internet's viewing happens, the structural floor is close enough that no more than one or two more big generational steps are likely on conventional metrics.

Two important caveats. First, that floor is not a single number — it depends on the distortion measure. On PSNR, the floor is close. On perceptual measures like VMAF or LPIPS, the floor is lower (you can be further from the original pixel by pixel and still look identical), and there is more room. Second, the RDP surface introduces a third axis — perception — and the floor there is harder to estimate because the "perceptual divergence" objective has many possible definitions, each with its own bound.

The practical implication is that the question "how much more compression is left" splits into three:

  • On PSNR and conventional MSE: small single-digit gains over VVC and AV2. The room is close to gone.
  • On perceptual measures like VMAF and LPIPS: 20 to 40 percent more room, mostly accessible only to neural codecs trained on those losses.
  • On the RDP perception axis: large remaining room, but the question shifts from "how few bits" to "how natural does the reconstruction look", which is a different objective altogether.

A 70 percent compression-improvement headline in 2026 is almost always measured on a perceptual axis that conventional codecs were never optimised for. That is real — neural codecs really do produce more natural-looking video at the same bitrate — but it is not the same number as the "60 percent better than HEVC" that AV1 measured against VMAF in 2018.

A horizontal bar chart titled BD-rate savings against H.264, generation by generation, averaged across JVET test sets at 1080p and 4K, where more negative is better. Six rows, each labelled on the left with the codec and its year and carrying its value inside the bar: H.264 (2003) has no bar and reads 0% (baseline); HEVC (2013) -50%; AV1 (2018) -65%; VVC (2020) -75%; AV2 (2025) -77%; DCVC-FM (2024) -82%. The bars are coloured by family: dark blue for the MPEG and ITU-T line, light blue for VVC as the frontier standard, green for the open AOMedia line, purple for the neural codec. A narrow dashed band near the right end of the scale is labelled Estimated Shannon floor: -90 to -92%. The axis runs from 0% to -100%. Notes add that DCVC-FM's -82% is measured on VMAF while the other five are PSNR-based BD-rate, and list the sources: Sullivan 2012, Chen 2020, Bross 2021, Li 2024, BBC R&D 2024, AOMedia 2025 Figure 3. Each generation closes part of the remaining gap. The dashed zone is a rough estimate of the Shannon floor on natural video at common streaming bitrates.

A common mistake: treating "Shannon limit" as a single number

The most expensive misconception in this area is the assumption that there is one Shannon limit you can quote for "video", the way there is one speed of light. There is not. The rate-distortion function R(D) depends on:

  • The source. A 4K sports stream has a different R(D) curve from a low-motion lecture video. Animated content has a curve different from both.
  • The distortion measure. PSNR, SSIM, VMAF, and LPIPS each yield different curves on the same source.
  • The statistical model. A simple block-IID Gaussian model gives one R(D); a temporally-correlated mixture model gives a much lower R(D); a deep autoregressive prior gives lower still.

So the right way to read a vendor claim is: "vs which codec, on which content, with which quality metric, and at which bitrate operating point?" A claim that reads "30 percent better than AV1" with no further context is approximately a coin flip — for some content and some metric it will be true, for some it will not. The Bjøntegaard methodology was invented to standardise this conversation, and the JVET common test conditions and AOM's CTC define exactly which sequences and which configurations are used in academic comparisons.16 If a vendor claim doesn't reference one of these test conditions, the claim is marketing, not measurement.

A worked example: the Shannon-limit budget for a 4K SDR catalogue

To anchor the numbers, imagine a streaming service with a 10,000-hour 4K SDR catalogue, encoded today in HEVC Main 10 at an average 12 Mbps. The catalogue weighs 10,000 × 3,600 × 12,000,000 ÷ 8 = 54 TB. CDN egress at 10x monthly catalogue plays is 540 TB.

Switching to AV1 at the same quality saves roughly 30 percent over HEVC on this kind of content: new bitrate around 8.4 Mbps, catalogue around 38 TB, egress around 380 TB. At a typical 2026 CDN price of about 0.01 dollars per GB on a tier-1 provider, that is (540 - 380) × 1,000 × 0.01 = 1,600 dollars per month, or 19,200 dollars per year saved on a relatively small library.

Switching to a future best conventional codec — AV2 or VVC — buys another 5 to 8 percentage points on top of AV1, perhaps another 1,000 dollars per month. Switching beyond that to a future neural codec trained on perceptual losses could buy another 15 to 20 percent, but only at the price of decoder cost — until consumer silicon ships a neural decoder, the savings are theoretical for delivery and only realised in cloud transcoding.

The arithmetic is the punchline. The remaining structural compression headroom on conventional codecs, on a real CDN budget, is a few thousand dollars per month per 10,000 hours of catalogue. That is enough to justify a codec swap every five or six years, but it is not the kind of number that funds a multi-year R&D program by itself. Neural codecs reset the math because they unlock a different axis (perception) and a different kind of saving, but they also reset the compute cost.

A pitfall to avoid: optimising the wrong objective

Engineers who already work in codecs sometimes spend optimisation budget driving down PSNR on operating points where PSNR no longer correlates with perceived quality. The classic symptom is a per-title encoding ladder that scores well on objective metrics but has waxy faces, smeared textures, and unnaturally smooth gradients on subjective review. The fix is to pick the metric (or the combination of metrics, including a perceptual one like VMAF and a learned-perceptual one like LPIPS) that matches your audience's actual judgement, and optimise for that. Otherwise the codec keeps moving down the R(D) curve on the wrong axis, and the picture looks worse at lower bitrate.

A two-by-two matrix titled Where each codec sits in 2026, placing the compute budget at the decoder against the quality objective. The columns are low compute (mobile, hardware decode) and high compute (GPU or cloud); the rows are perceptually natural and pixel accurate. Top left, Future hybrid codecs: perceptual targets inside conventional pipelines; when to ship, not yet, watch AV2 extensions and neural hybrids from 2027. Top right, Neural codecs: DCVC-FM, DCVC-RT and MS-ILLM match or beat VVC on VMAF; when to ship, cloud pre-encoding for VOD, not yet a delivery codec. Bottom left, Mainstream delivery: HEVC and AV1, broad hardware decode, low watts; when to ship, the default for OTT and live in 2026. Bottom right, Frontier conventional: VVC, AV2 and JVET ECM, best on PSNR and slow to encode; when to ship, premium tiers and broadcast. A closing line says codec choice in 2026 is a two-axis decision: the compute your decoder has, and what your audience actually judges quality by Figure 4. The same decision as a map: what your decoder can afford, and which objective your audience actually reacts to.

Where Fora Soft fits in

We build video-streaming, OTT, e-learning, telemedicine, video-conferencing, and AR/VR systems where codec-efficiency decisions show up in the line items every quarter. The framing we apply to product teams is the same one this article ends on: separate the question "how few bits do we need for our content at our quality bar" from the question "how few bits is possible in principle". The first answer is your engineering target; the second tells you when to stop expecting more from each generation. In OTT and e-learning deployments we routinely measure first-party R(D) curves on customer content rather than rely on standardised test sets, because the wrong baseline assumption can mis-size a CDN contract by 20 percent in either direction.

Video Encoding: A Complete Guide to Video Compression - book cover

The book

Video Encoding: A Complete Guide to Video Compression

This article draws the ceiling; the book measures the distance to it. Rate-distortion theory is worked through end to end — what R(D) is for a real source, how the curve moves when you change the distortion measure, and how a BD-rate number is computed, so you can check a vendor's claim against the test conditions it came from. The codec chapters then walk the same ladder this article plots: what H.264, HEVC, AV1, VVC and AV2 each bought, where the tools ran out, and what learned compression optimises instead. Written by Nikolay Sapunov, CEO at Fora Soft.

Read it on Amazon →Kindle & paperback · Fora Soft Video Engineering Handbook

How much compression headroom is actually left on your catalogue?

Send us your current ladder and a handful of representative titles, and we will measure the rate-distortion curve on your own content: what a codec swap is worth at your quality bar, which metric your audience reacts to, and what the change does to your CDN bill and your encoder farm. You get numbers from your library rather than from a JVET test set. 250+ video products shipped since 2005.

Book a 30-min call →WhatsApp →Email us →

See our case studies


  1. Shannon, C. E. (1948). "A Mathematical Theory of Communication." Bell System Technical Journal. The founding paper of information theory; defines entropy and proves the source coding theorem. https://people.math.harvard.edu/~ctm/home/text/others/shannon/entropy/entropy.pdf. Accessed 2026-05-16. ↩

  2. Shannon's source coding theorem — Wikipedia overview of the lossless compression bound at the entropy rate. https://en.wikipedia.org/wiki/Shannon%27s_source_coding_theorem. Accessed 2026-05-16. ↩

  3. Rate-distortion theory — Wikipedia. Definition of R(D), the smallest rate at which a source can be encoded subject to a distortion constraint. https://en.wikipedia.org/wiki/Rate%E2%80%93distortion_theory. Accessed 2026-05-16. ↩

  4. Bjøntegaard, G. (2001). "Calculation of average PSNR differences between RD-curves." VCEG-M33. The Bjøntegaard delta methodology that codec teams use to compare two codecs at the same quality. https://www.itu.int/wftp3/av-arch/video-site/0104_Aus/VCEG-M33.doc. Accessed 2026-05-16. ↩

  5. Sullivan, G. J. et al. (2012). "Overview of the High Efficiency Video Coding (HEVC) Standard." IEEE Transactions on Circuits and Systems for Video Technology. Reports ~50% BD-rate savings of HEVC over H.264 on JVET test sets. https://ieeexplore.ieee.org/document/6316136. Accessed 2026-05-16. ↩

  6. Chen, Y. et al. (2020). "An Overview of Coding Tools in AV1: the First Video Codec from the Alliance for Open Media." APSIPA Transactions on Signal and Information Processing. AV1 averages ~30% BD-rate gain over HEVC on random-access configurations. https://www.cambridge.org/core/journals/apsipa-transactions-on-signal-and-information-processing/article/overview-of-coding-tools-in-av1-the-first-video-codec-from-the-alliance-for-open-media/. Accessed 2026-05-16. ↩

  7. Bross, B. et al. (2021). "Overview of the Versatile Video Coding (VVC) Standard and its Applications." IEEE Transactions on Circuits and Systems for Video Technology. VVC reports ~40-50% BD-rate gain over HEVC on UHD content. https://ieeexplore.ieee.org/document/9503377. Accessed 2026-05-16. ↩

  8. Performance Comparison of VVC, AV1, HEVC, and AVC for High Resolutions (2024). MDPI Electronics. Reports BD-rate savings at 8K: VVC -78%, AV1 -63%, HEVC -53% vs H.264. https://www.mdpi.com/2079-9292/13/5/953. Accessed 2026-05-16. ↩

  9. AOMedia AV2 open video codec release nears, delivers around 40% bandwidth reduction (2025-11). CNX Software. AOMedia's reported AV2 efficiency over AV1; YUV-PSNR 28.6%, VMAF 32.6% gains. https://www.cnx-software.com/2025/11/21/aomedia-av2-open-video-codec-release-nears-delivers-around-40-bandwidth-reduction/. Accessed 2026-05-16. ↩

  10. Blau, Y., Michaeli, T. (2018). "The Perception-Distortion Tradeoff." CVPR 2018. The original paper proving that distortion and perceptual quality are mathematically at odds. https://arxiv.org/abs/1711.06077. Accessed 2026-05-16. ↩

  11. Blau, Y., Michaeli, T. (2019). "Rethinking Lossy Compression: The Rate-Distortion-Perception Tradeoff." ICML 2019. Introduces the RDP function as a generalisation of Shannon's R(D). https://arxiv.org/abs/1901.07821. Accessed 2026-05-16. ↩

  12. Li, J., Li, B., Lu, Y. (2024). "Neural Video Compression with Feature Modulation." CVPR 2024. DCVC-FM beats VTM (the VVC reference) by 25.5% in BD-rate under intra-period -1. https://arxiv.org/abs/2402.17414. Accessed 2026-05-16. ↩

  13. Jia, Z. et al. (2025). "Towards Practical Real-Time Neural Video Compression." CVPR 2025. DCVC-RT delivers ~21% BD-rate gain over VVC at real-time framerates. https://dcvccodec.github.io/. Accessed 2026-05-16. ↩

  14. Benchmarking Conventional and Learned Video Codecs (2024). arXiv preprint. Cross-codec comparison of DCVC-FM, ECM, AVM, and AV1 on VMAF. https://arxiv.org/abs/2408.05042. Accessed 2026-05-16. ↩

  15. Liu, J., Lu, G., Wang, Y. (2025). "Advances in Neural Video Compression: A Review and Benchmarking." Preprints.org. Reviews how close current codecs sit to the theoretical R(D) bound on natural video. https://www.preprints.org/manuscript/202604.0035. Accessed 2026-05-16. ↩

  16. JVET Common Test Conditions and AOMedia Common Test Conditions — the standardised test sets and configurations used to make BD-rate comparisons meaningful. https://www.itu.int/en/ITU-T/studygroups/2017-2020/16/Pages/video/jvet.aspx. Accessed 2026-05-16. ↩