Travel Photography

How a Rejected NSF Proposal Built the Invisible Foundation of the Digital World

In early 1972, Kansas State University professor Nasir Ahmed submitted a research proposal to the National Science Foundation seeking financial backing to study a novel cosine transform utilizing Chebyshev polynomials. The NSF rejected the application. According to Ahmed’s later recollections, a reviewer dismissed the premise as "too simple" to warrant public funding. More than five decades later, that exact mathematical framework underpins the arithmetic of virtually every Joint Photographic Experts Group (JPEG) image file generated, stored, transmitted, and displayed worldwide.

The historical trajectory of this foundational signal processing technology illustrates a persistent challenge in scientific innovation: peer-review systems designed to filter out unpromising concepts occasionally fail to recognize paradigm-shifting simplicity. Today, Ahmed’s discrete cosine transform (DCT) is acknowledged as one of the most commercially and culturally impactful mathematical formulas of the modern digital age, having quietly powered the expansion of the internet, digital photography, and global telecommunications.

The Academic Environment and the Search for a Stand-In

During the late 1960s and early 1970s, the field of image coding was experiencing rapid, chaotic growth. Academic institutions, including the University of Southern California’s Image Processing Institute and UCLA, were flooded with newly introduced orthogonal transforms. Researchers evaluated these competing mathematical models through subjective, qualitative assessments—manually inspecting test images that had been artificially compressed and reconstructed.

As the discipline moved toward rigorous measurement, the Karhunen-Loeve transform (KLT) emerged as the gold standard. Under a first-order Markov model of an image, the KLT is mathematically proven to achieve the lowest possible mean-square error. However, the transform suffered from a crippling practical limitation: because it depends entirely on the specific statistical properties of the exact data input, engineers could not design a fast, fixed algorithm to compute it efficiently.

Seeking an alternative, Ahmed scoured academic literature for a signal-independent transform that could approximate the performance of the KLT while offering high computational speed across any arbitrary block of pixels. His breakthrough came from a 1968 Prentice-Hall textbook detailing Chebyshev interpolation for computer evaluations. He discovered that specific cosine functions closely mirrored the basis functions of the Karhunen-Loeve transform across the precise range of pixel-to-pixel correlations found in standard photographs.

Persistence Through Rejection and the Summer of 1973

When the NSF declined Ahmed’s funding request, the academic setback did not halt his research. Operating without external grants, Ahmed redirected his focus, bringing in his doctoral student T. Natarajan and his colleague K. R. Rao from the University of Texas at Arlington.

The team dedicated the summer of 1973 entirely to refining and testing the transform. The initial empirical results were so high-performing that Ahmed initially suspected an error. Seeking external validation, he approached Harry Andrews, a leading researcher from the USC group, during a conference session on Walsh functions in New Orleans. Andrews mailed Ahmed a computer program to evaluate the transform against the stringent rate-distortion criterion.

The results confirmed Ahmed’s hypothesis. The discrete cosine transform performed nearly as well as the theoretical ceiling of the KLT while outperforming all alternative models under consideration. Andrews urged Ahmed to publish immediately.

In January 1974, Ahmed, Natarajan, and Rao published their findings in IEEE Transactions on Computers. To secure rapid publication, they submitted the work as a brief correspondence item rather than a full-length manuscript, ensuring it appeared across pages 90 through 93 of that month’s issue.

Mechanics of the Discrete Cosine Transform in Modern Image Compression

To understand the ubiquity of Ahmed’s mathematics, one must examine how a standard digital encoder processes a color photograph. When an image is saved as a JPEG, the encoder converts the visual data into a luminance (brightness) channel and two chrominance (color) channels. Grayscale files utilize a single channel, while CMYK formats utilize four, but the underlying tiling process remains identical.

The encoder divides each image channel into discrete tiles measuring eight pixels by eight pixels, yielding blocks of 64 individual numerical values. Before applying the transform, the encoder subtracts 128 from each 8-bit sample, centering the data range between -128 and 127.

The two-dimensional discrete cosine transform is then executed across the 64-number tile. Sixty-four values enter the algorithm, and 64 transform coefficients emerge. Crucially, these output numbers do not represent raw pixels; rather, each coefficient represents the statistical weight of a fixed spatial pattern.

The top-left coefficient—known as the DC coefficient—represents eight times the average value of the level-shifted tile. The remaining 63 AC coefficients denote the precise presence of varying spatial frequencies and directional details within the tile. Because real-world photographic scenes, such as human skin, blue skies, or soft shadows, exhibit high spatial continuity, their tiles concentrate nearly all visual energy into the top-left corner, leaving fine high-frequency patterns close to zero.

At this stage, the transformation remains mathematically reversible down to minor rounding errors; no actual data has been discarded. The adoption of the 8×8 block size itself emerged from international standardization efforts in the late 1980s. The JPEG committee formally selected the 8×8 adaptive DCT proposal at a meeting in Copenhagen in January 1988, finding the dimension struck an optimal balance between low computational overhead on 1980s hardware and adequate compression efficiency.

Quantization, Artifacts, and the Quality Slider

Data loss occurs exclusively during the subsequent quantization phase. Each of the 64 transform coefficients is divided by a corresponding integer specified within an 8×8 quantization table, and the resulting quotient is rounded to the nearest whole number.

The design of these tables exploits human psychovisual limitations. Because the human eye tolerates significantly more visual error in fine spatial textures than in overall luminance variations, quantization tables apply heavier divisors to high-frequency coefficients.

When a user adjusts a software application’s "quality" slider, the program scales these divisors dynamically. At lower quality settings, aggressive division turns high-frequency coefficients into zero. During the final encoding stage, the algorithm reads the 64 coefficients in a zigzag pattern from lowest to highest frequency, clumping the accumulated zeros together where run-length and Huffman coding reduce their file size footprint to near zero.

This mathematical pruning explains the common visual failure modes of aggressive JPEG compression. When divisors are excessively coarse, adjacent tiles reconstruct to slightly different average values, revealing the underlying eight-pixel grid—a phenomenon commonly observed as "banding" in compressed digital skies. Similarly, sharp edges develop faint halos because the fine spatial patterns required to render crisp boundaries have been zeroed out.

Legacy, Recognition, and Continued Relevance

For decades, the intellectual authorship of the discrete cosine transform remained obscured by its foundational nature. Because Fourier analysis and Chebyshev polynomials are public mathematical domains, the specific applied formulation developed by Ahmed, Natarajan, and Rao was frequently treated as an unauthored natural resource.

However, historical records clearly distinguish Ahmed’s specific operational framework from general trigonometric principles. His deliberate decision to publish the methodology in 1974 established a public technical standard that the emerging JPEG committee could readily adopt in 1986 without proprietary licensing barriers.

Public recognition for Ahmed accumulated slowly over the decades. Following his relocation to the University of New Mexico in 1983, where he later served as electrical and computer engineering department chair and interim dean of engineering, his contributions remained largely confined to academic and engineering circles. Cultural mainstream awareness arrived unexpectedly in February 2021, when NBC’s drama series This Is Us dramatized his struggle and subsequent breakthrough during an era when global populations relied heavily on video conferencing technologies powered by his mathematics.

Official institutional honors followed in subsequent years. In 2026, the National Academy of Engineering elected Ahmed as a member, and the Institute of Electrical and Electronics Engineers (IEEE) presented him with the prestigious Fourier Award for Signal Processing. The award citation explicitly recognized his foundational contributions to the digital revolution through the development of the discrete cosine transform—arriving 54 years after an NSF review panel concluded his proposal was "too simple."

Despite decades of subsequent computational research, newer digital image formats continue to build upon Ahmed’s original framework rather than discarding it. High Efficiency Video Coding (HEVC)—utilized for Apple’s HEIC image format—and JPEG XL’s VarDCT mode both employ integer approximations of discrete cosine transforms across scaled block sizes ranging from 2×2 up to 256×256. More than half a century after its initial rejection by federal grant reviewers, the modern visual landscape remains fundamentally defined by an overlooked 1972 proposal for a simpler way to compute cosines.

Related Articles

Leave a Reply

Your email address will not be published. Required fields are marked *

Back to top button