Online Academy for Audio Engineering & Music Production

Immersive Remix – How Does Stereo Become 3D?

At least since Apple Music added the “Spatial” category to its program, immersive music has been on everyone’s lips… or rather, on everyone’s AirPods, earbuds, soundbars, etc. In particular, the available catalogue of Dolby Atmos mixes is growing rapidly, and statistics show that listeners often prefer immersive audio formats to traditional stereo sound. Other streaming services such as Amazon Music and Tidal also offer music in 3D.

By now, not only major-label artists but also more and more studios and engineers have upgraded their systems. A lot has happened in the field of DAW manufacturers and DIY distributors as well. As a result, the path to Dolby Atmos, binaural audio, Ambisonics, Sony 360 Reality Audio, Auro-3D and other immersive audio formats is now open to indie and DIY artists as well.

This article will give you an introduction to remix in immersive audio formats so that your stereo production can really shine in 3D. Let’s get started!

Preparation is everything

As with any mix, it’s the ingredients that determine the result – so we need good signals first! We generate these from the original stereo mix. Anyone who works with hip-hop, electronic, rock or pop music knows how essential the use of effects can be for the sound design of a song. But even when working with jazz or classical music, it makes sense to preserve as much as possible of the original sonic vision of the stereo production rather than completely reinventing the piece. It takes sensitivity – on the one hand, we want to make the best possible use of the immersive space, while on the other hand preserving the artistic intention as faithfully as possible.

That’s why we don’t start with individual tracks, but with stems – grouped signals that include the original effects. What sounds great in stereo usually sounds great in 3D as well. 3D mixes generally require a few more stems than you might use for stem mastering – 12–24 stereo signals is a good guideline for an average Dolby Atmos pop mix.

In practice, it has proven useful to keep all effects and additionally bounce reverbs and echoes as separate stereo stems. In the 3D mix, these spatial elements can then be moved toward the rear or overhead, duplicated or varied according to taste, saving you the trouble of recreating the original reverb space by ear in 3D. Another option is to continue working directly in the original DAW session and, for example, replace a stereo reverb with a 7.1.2 reverb.

Screenshot eines Datei-Ordners, der 17 Stems für 3D-Remix und das Master enthält.
Possible stems for a 3D mix

De-mixing instead of upmixing

A classic scenario: The hard drive containing the original session is gone, or the project simply no longer exists – all that remains is the finished stereo mixdown. So how can you still create a good 3D mix? We’ll spare you the backup lecture at this point 🙂

The obvious solution is to place an upmix plugin on the finished stereo mix and voilà: 3D sound. Temptingly simple – but not always a good idea. An upmix plugin can only analyze roughly where something is positioned within the stereo image and distribute the sound across additional speakers based on that estimate. It cannot distinguish what is actually playing – vocals, bass and drums remain one inseparable whole. The result is often a diffuse cloud of sound rather than a mix in which every element has its own clearly defined position in the room.

The better alternative is called de-mixing, also known as stem separation. The idea is this: Instead of simply spreading the stereo mix wider, you try to transform it back into its original components – for example, creating separate tracks for vocals, bass, drums and so on. In a way, you are “undoing” the mix.

The problem is that once several instruments have been mixed together, their frequencies overlap. Cleanly separating those overlaps again is extremely difficult, and this is exactly where artificial intelligence comes into play. Modern AI models have been trained on vast amounts of music and have learned to recognize typical sound patterns of vocals, drums, bass and other elements even when they overlap. The result is never perfect, but it is often surprisingly usable.

Screenshot einer Ordnerstruktur. Die oberste Ordnerebene enthält den Stereo-Song, im Ordner befinden sich 4 extrahierte Stems.
Extracted stems

There is now a wide range of software available that can separate your song into stems, such as Audionamix Xtrax or iZotope RX Music Rebalance. Logic Pro X also has this feature on board. There are also online services such as lalal.ai, where the separation is performed remotely and no software needs to be installed on your own system.

De-Mixing mit Izotope RX Music Rebalance
De-Mixing with iZotope RX Music Rebalance

In these HOFA College videos, we take a closer look at AI stem splitting and compare established tools.

If necessary, the separated stems can then be manually processed or restored before creating a new mix from them. You can, of course, still use upmix plugins on individual stems, or play the signals through speakers and record them in 3D in a suitable room using a large microphone setup to create multichannel audio. You can also use the stereo mix as a foundation and simply supplement it with some of the separated stems – the possibilities are endless.

With the HOFA online course 3D Audio, you will learn everything you need to create successful immersive audio productions.

Save 10% in the introductory period

Routing & templates

Now let’s look at a few ideas you may encounter when taking your first steps into the world of immersive mixing. Since Dolby Atmos is currently the format with the greatest demand in music production, we’ll focus on Atmos here.

If an original stereo version of the song already exists, it is advisable to use the stereo master as the starting reference for the 3D mix. After all, the immersive version will ultimately be compared with the master on streaming platforms, and mastering often introduces substantial changes to the overall sound.

The stereo master also serves as a reference for the export length of the Atmos master. Lead-ins, fades, arrangement and all other timing-related elements must match exactly between the two formats, and the files need to be precisely in sync, with a maximum deviation of 50 ms, for the Atmos mix to be accepted by digital streaming platforms. To make it easy to switch between Atmos and stereo while comparing, the stereo reference is included directly in the Dolby Atmos mixing session.

Der Screenshot zeigt den Stereo Referenz Track in der ProTools Session. Er ist auf ein Objektpaar namens "StereoThrough" geroutet.
Stereo reference track

Instead of a stereo mix bus, Dolby Atmos uses a 10-channel bed bus and 118 mono sums for the 118 objects. Whether SSL, Neve, API, Fairchild, Manley or custom-built – many rock and pop engineers swear by their bus compressor and naturally don’t want to do without it in 3D.

The only problem is: There is no mix bus!

However, bus processing can still be recreated through a few workarounds. This requires multiple plugin instances placed on the individual channels. If you want to control all of these plugins from a single instance, their parameters can be linked.

Compression can also benefit from sidechaining. For example, a stereo master can be used as a sidechain signal so that compressors or limiters on the individual channels are all controlled by the same trigger signal. With a bit of routing and some manual setup, workflows like this can be recreated even without a conventional stereo mix bus.

Der Screenshot von ProTools zeigt die SSL Bus Compressor Instanz auf dem Bed Master. Auf der rechten Seite sind die Gruppeneinstellungen für alle Masters zu sehen. Die Inserts sind gelinkt.
Master dynamics with sidechain input and group link

The LFE channel is unfortunately often misunderstood by music creators. It is not a subwoofer channel that is supplied with the low frequencies of the other speaker channels by means of a crossover. Instead, it is a completely independent channel that can receive the full frequency spectrum.

Der Screenshot von ProTools zeigt die Inserts eines LFE Kanals: Ein Subharmonic Plugin und einen Low Pass bei 120 Hz.
LFE inserts

Of course, we also need some send effects when we want to turn stereo into 3D. It is often quite easy to use the original stereo reverbs or echoes in 3D by duplicating them two or three times and placing them in the 3D space with slightly different settings. Often, simply delaying the rear spatial components slightly is enough.

Alternatively, there are now several reverb and delay plugins designed specifically for 3D. However, you should be rather careful with multichannel reverbs in Spatial Audio. Not everything that sounds great on speakers will translate well to headphones! The more reverb channels there are, the greater the chance that the reverb tails will build up during headphone rendering, resulting in an overly “wet” headphone mix. If you use 3D reverb, it is therefore advisable to keep its level relatively low.

Der Screenshot von ProTools zeigt 3 Duplikate eines Stereo-Reverb-Kanals, die mit dem Plugin "DMG TrackControl" um 40, 60 bzw. 100 ms verzögert wurden.
Duplicated and delayed stereo reverb stems

Mastering & export

Delivering a Dolby Atmos mix comes with a number of specifications that you should keep in mind throughout the process.

Loudness: Why -18 LUFS?

Dolby Atmos specifies a maximum loudness of -18 LUFS and a maximum true peak of -1 dBTP. If those terms are new to you, here’s a quick explanation:

  1. LUFS measures how loud a song is perceived on average, rather than simply measuring its peak level.
  2. dBTP (True Peak) describes the highest actual signal peaks, including peaks that can occur between digital samples and might otherwise be overlooked.

The infamous “Loudness War” – the constant race to make mixes louder and increasingly flattened through compression – simply doesn’t happen in Dolby Atmos because the specification prevents it. You can, of course, still compress your mix beyond good taste, but there is a clearly defined upper limit to its loudness. Loudness metering is built directly into the Dolby Atmos Renderer, so there is no need for an external tool.

Der Screenshot vom Dolby Atmos Renderer zeigt das Loudness Analyse Tool.
Loudness analysis with the Dolby Atmos Renderer

If you work for a label, you will usually receive a technical delivery specification document containing precise requirements for loudness, file formats, channel assignments and sometimes even creative aspects, such as the desired degree of spatiality or object usage. You should read these specifications carefully and follow them before delivering your master.

Final polishing with compressors, EQs or saturation is common. However, unlike in the stereo world, this processing is not applied to a conventional mix bus but on an object or bed basis, because there is no traditional stereo sum in Dolby Atmos.

Translation to different playback systems has always been an important aspect of mastering, but with object-based formats it has taken on a completely new significance. Since the format is interpreted individually for almost every conceivable speaker configuration and can even be rendered binaurally for headphones, the differences between playback systems can be enormous. It takes experience to find the sweet spot that works across all of them.

Conclusion

There are many ways to turn a stereo mix into a convincing 3D remix. Well-prepared stems make the process particularly straightforward, but de-mixing can also turn an existing production into the foundation for an immersive mix. From there, the key factors are thoughtful routing, the right approach to spatial effects and reverbs, and a master that translates successfully across different playback systems.

Want to dive deeper and create professional immersive productions yourself? In the new HOFA-College online course 3D Audio, you’ll learn the complete workflow for Dolby Atmos, 3D recording, mixing and mastering – with a practical approach that includes creating your own immersive mix.

The “3D Audio” online course is also included at a discount of more than 50% in the HOFA AUDIO DIPLOMA distance-learning bundle. More information ›

Author

Picture of Christoph Thiers
Christoph Thiers
Christoph Thiers has been active in the music industry for over a decade and has worked on hundreds of productions of various genres as recording, mixing and mastering engineer. His track record includes artists such as Die Fantastischen Vier, Sarah Connor, Birdy, Nathan Evans, RAF Camora and Boris Brejcha, as well as numerous awards and chart placements. He is also engaged in new media formats and artist development, acts as a consultant to indie labels, artists and start-ups alike and has been involved in various software developments for professional music production. In recent years, Christoph has specialised in immersive music production and handles Dolby Atmos mixes for international label clients and renowned indie artists.

One Response

  1. Nice work! I’m an algorithm engineer and trying to find a common solution to remix stereo pop music into 3D, my confusion is where to put the guitar, bass and vocal in the virtual 3D room. I’m not pro in remixing so it’s difficult to me. Are there any suggestions ?

Leave a Reply

Your email address will not be published. Required fields are marked *

Do you have any questions about the HOFA audio engineering courses?