Two-part positioning: Part 1 reflects DSD's real limitations in the MASH modulator era of low compute power — about 100 misconceptions pointing at that era's DSD. This part is based on the DpdoEngine v6.59+ toolchain — high-order (9th) modulator + Kakeya polynomial compression + long-tap FIR (65536-131072 taps) + all-platform SIMD acceleration — discussing DSD's real face today.
Preface: Why a Part 2 Is Needed
A technology's face is determined by two things: theoretical foundation and implementation quality.
DSD's theoretical foundation (1-bit ΔΣ modulation + noise shaping) was laid in 1962; the SACD standard shipped in 1996. In the 25 years since, most technical criticism of DSD targeted not its mathematics but its implementation quality — low-order MASH modulators, finite-tap FIRs, noise shaping held back by compute cost, poor time-domain precision, messy distortion spectra.
Since DpdoEngine v6.59, three technical lines arrived together, making "high-quality DSD on consumer hardware" real:
- High-order single-loop modulator (9th order, default): noise-shaping slope ~110dB/oct (over 7× the ~15dB/oct of MASH 3-4 order equivalents)
- Kakeya polynomial coefficient compression: lets 65536-131072-tap FIRs run real-time on consumer CPUs, solving the compute bottleneck of long taps + high-order modulator stability
- AVX-512/AVX2/NEON SIMD all-platform acceleration: makes the above computation commercially viable in real time
These are not "optimizations" — they are enabling technologies. Without them, consumer hardware cannot run high-quality DSD conversion.
Every technical criticism in Part 1 assumed "MASH + compute bottleneck = the whole picture of DSD." Today, that premise is no longer the whole truth.
Q1: DSD64's noise shaping is too poor, inferior to DSD256?
Misconception: DSD64's in-band SNR is only around 110dB; you need DSD128/256 for usable noise shaping. So DSD64 is doomed to low-end scenarios.
Truth (current): That was a MASH-era number. Under DpdoEngine v6.60 default config (order=9, taps=65536, Kakeya 128th-order compression), DSD64's 20Hz-20kHz in-band SNR (A-weighted) measures about 125dB — exceeding MASH-era DSD256 specs. No visible idle-tone spikes in silence; harmonic distortion baseline drops from ~-95dBFS to ~-108dBFS.
Previously it took four times the rate to compensate for noise-shaping shortfalls. Now DSD64's own noise shaping is sufficient. High rates have gone from necessity to option — meaningful only when a DAC architecture clearly benefits from higher rates.
Q2: DSD's time-domain performance is poor; impulse response is blurred?
Misconception: DSD's impulse response is not good enough (jittery, scattered, blurred); dynamic transient detail is worse than PCM.
Truth (current): True in the MASH era — low-order shapers produced heavy pulse jitter in the time domain, transients drowned in noise fluctuation. The long-tap FIR + high-order modulator combination changes this:
- Step response overshoot: MASH ~12% → current ~3% (65536 taps, Kaiser window)
- Group delay variation 20Hz-20kHz: ±5 samples → ±0.5 samples
- Impulse response tail (after -60dB): ~100μs → ~30μs
Time-domain precision improvements map directly to what listeners call "transient clarity." The old criticism of DSD blurriness no longer holds under current implementation.
Q3: DSD's "analog warmth" is harmonic distortion?
Misconception: DSD's harmonic distortion is clearly higher than PCM; the so-called "warmth" is distortion misheard as character.
Truth (current): The Part 1 data (19th harmonic at -105dBFS) came from v6.55 and earlier MASH topologies. v6.60's 9th-order modulator (optional 11th) with optimized noise shaping pushes all high-order harmonics down 12-18dB:
| Harmonic | MASH (v6.55-) | v6.60 | Improvement |
|---|---|---|---|
| 3rd | -92dBFS | -104dBFS | 12dB |
| 5th | -96dBFS | -112dBFS | 16dB |
| 7th | -98dBFS | -114dBFS | 16dB |
| 19th | -105dBFS | -122dBFS | 17dB |
Current DSD harmonic distortion spectra are near the noise floor (~-108dBFS). DSD's "warmth" can no longer be explained by harmonic coloration — a more accurate description is the influence of the noise shaper's spectral distribution on auditory masking. When you need to look below -108dBFS for differences, most systems' noise floors have already covered it.
Q4: Upsampling to DSD256/512 is always better than DSD64?
Misconception: Higher rates are always better; DSD64 isn't worth using; DSD256/512 is the ultimate.
Truth (current): When DSD64 is already good enough, the marginal benefit of high rates diminishes sharply:
| Comparison | SNR gain | Compute | File size |
|---|---|---|---|
| DSD64→128 | ~6dB | 2× | 2× |
| DSD128→256 | ~3-4dB | 2× | 2× |
| DSD256→512 | ~1-2dB | 2× | 2× |
Recommended strategy:
- Default DSD64. ~123dB in-band SNR is enough for 99% of systems and listening environments
- Go DSD128 when: the DAC shows significantly better measured THD+N and dynamic range in high-rate DSD mode
- DSD256/512 not recommended as daily config unless your DAC is specifically optimized and transparent
In blind comparisons, DSD64 vs DSD128 is no longer reliably distinguishable on most systems — a noise-shaping difference, not a listening-level one.
Q5: DSD file sizes are too big, impractical?
Misconception: DSD64 is 2-3× the size of FLAC CD, fundamentally unsuitable for daily use.
Truth (current): File size is indeed larger than FLAC, but understand what it's doing. A large portion of DSD64's 2.8Mbps bitrate is not "audio information" — it's bandwidth consumed by 1-bit encoding itself. 1-bit means each sample carries 1 bit, requiring extremely high sample rates (2.8MHz) to sustain information content. This is an encoding-efficiency trade-off, not "DSD wasting space."
Two facts:
- DSD64 is about 2× the size of CD FLAC — no practical obstacle given today's storage costs and bandwidth (100MB-scale files, TB-scale drives)
- With DST compression, DSD64 can compress to ~1.4-1.5Mbps, close to FLAC CD
The "files too big" criticism was a real pain point in DSD's early days (2000s: 750MB CD-Rs, 40GB hard drives). Looking back from 2026, this is not a technical problem.
Q6: DSD's high-frequency extension is fake — just a noise layer?
Misconception: DSD's "high-frequency extension" up to 50-100kHz is entirely noise-layer artifacts, with no substance.
Truth (current): The statement itself is correct — DSD energy above 20kHz is mainly noise-shaping residue, not musical signal. But between "it's all noise layer" and "it's meaningless" there's a layer being missed.
DSD noise shaping's core operation pushes noise from the audible band to high frequencies without changing total noise energy. This means:
- Noise in 20Hz-20kHz genuinely decreases — that's the purpose, and it succeeds. DSD64's 123dB in-band SNR is the product of this process
- There is indeed a noise layer at 50-100kHz, but its spectral distribution and amplitude are controllable through modulator design and FIR parameters. A good modulator gives this noise an approximately "pink-blue" distribution, minimizing correlation with the signal
- The DAC's low-pass filter quality determines how much this out-of-band noise is attenuated — so DSD playback depends heavily on DAC design. Not DSD's fault; it's a system design matter
"The high-frequency extension is fake" is indeed correct. But what matters is DSD's SNR in the audible band, not the controllable noise layer above 50kHz.
Q7: DSD's editing inconvenience is a format flaw?
Misconception: DSD can't be edited directly, so it's unsuitable for any serious audio workflow.
Truth (current): This hasn't changed, and doesn't need to. DSD doesn't support linear operations (EQ, dynamics, reverb, editing) in the 1-bit domain — that's determined by the information structure of 1-bit PDM encoding. Today, the PCM→DSD conversion + SACD ISO authoring workflow in DpdoEngine is mature enough; DSD plays the role of final output format, not production format:
PCM recording → editing/mixing/mastering (PCM domain) → high-quality upsampling to DSD (DpdoEngine) → SACD ISO authoring (DpdoPackSACD) / DSF archive
This is a mature "produce in PCM, deliver in DSD" division. Streaming and portable devices go PCM; hi-fi playback and SACD releases go DSD. The two are not substitutes — they collaborate.
Q8: High-order modulators + long FIRs are just icing on the cake?
Misconception: The upgrade from MASH to 9th-order + 65536-tap FIR is "better DSD" — no real impact on ordinary users.
Truth (current): This is the point in this part most needing to be understood.
Why couldn't MASH-era DSD push noise shaping high? Because low-order modulators have shallow noise-shaping slopes; pushing DSD64's in-band noise low enough required increasing OSR — and doubling OSR doubles compute. On then-current hardware, DSD128 was the ceiling; going higher was impossible.
Why don't high-order modulators need OSR doubling? Because the ~110dB/oct noise-shaping slope means 110dB of noise attenuation per octave — at the same DSD64 OSR (64×), a 9th-order modulator achieves about 12dB better in-band SNR than 3rd-order MASH. This gap comes directly from noise-shaping efficiency, independent of rate.
Kakeya compression exists because 65536-tap FIR compute on AVX-512 reaches 14μs per sample, while DSD64's real-time window is only 0.35μs — over 40× over budget. Without Kakeya compressing FIR coefficients into polynomials (128th order, 1/512 compression), long FIRs can't run. It's not "optimized well" — it's "without it, this can't be done at all."
So "from MASH to 9th-order + long FIR" is not incremental improvement — it's a qualitative enabling leap. It doesn't change DSD's theoretical framework — it lets DSD's theoretical potential be nearly realized on consumer hardware for the first time.
Q9: Pursuing DSD requires a top-tier system?
Misconception: DSD only matters on systems worth hundreds of thousands; ordinary audiophiles shouldn't touch it.
Truth (current): The real threshold isn't "how expensive the system is" but "how good the DAC's DSD mode is." A rule of thumb:
Open your DAC's spec sheet (or find measured data) and check:
- THD+N and dynamic range in DSD mode
- Corresponding figures in PCM mode
If DSD mode beats PCM mode in THD+N and dynamic range by >3dB, then DSD upsampling on your system can produce an actually audible difference. This is unrelated to system price; it's about the quality gap between the DAC chip's DSD and PCM paths.
The most common consumer benefit scenario: a DAC with an ESS/Rivet/AK445x chip, where THD+N in DSD mode is typically 3-5dB better than PCM mode. Paired with a desktop speaker/headphone system with a noise floor below 25dBA, upsampling to DSD64 can create a stable audible difference.
The threshold isn't price — it's whether DAC, backend and environment match up.
Q10: Will all music become DSD in the future?
Misconception: After DSD technology improves, it will become the mainstream streaming format.
Truth (current): No. DSD replacing PCM in the mainstream consumer market is a zero-probability event — the reasons were laid out in Part 1: streaming doesn't support it, editing is unfriendly, hardware coverage is low.
But that's a separate question from "does DSD have value." Something doesn't need to become mainstream to justify its existence. SACD releases, hi-fi playback and recording archives — these scenarios need neither streaming coverage nor 2 billion users. They need a good toolchain — which was missing before high-order modulators + long FIRs, making DSD not good enough in those scenarios either. Now that the toolchain is in place, the "not good enough" obstacle is cleared.
DSD's reasonable niche today: high-fidelity playback with high-quality DSD upsampling on specific DAC architectures in quiet listening environments, plus SACD-format release and archiving.
It needs no more. It never did.
Conclusion: Part 1 and Part 2 Belong Together
The 130 items in Part 1 dissected MASH-era DSD limitations. Those criticisms each held under then-current implementation. Part 2 says something simple:
When the toolchain is good enough — high-order modulator + long-tap FIR + Kakeya compression + SIMD acceleration — DSD is no longer the DSD Part 1 described. Many "innate flaws" are actually "implementation flaws."
This is not saying DSD beats PCM. A rational user should:
- Read Part 1 to understand what MASH-era DSD did and why it was criticized
- Read Part 2 to understand what the current toolchain changed and what remains unchanged
- Then decide whether to use DSD in their system based on DAC architecture → backend → listening environment → program material
Not good because of Kakeya. Not good because of 9th order. Good because your system happens to match on this chain.