THUNDERBOOM

MUSIC TECH

INNOVATION LAB

THUNDERBOOM

MUSIC TECH

INNOVATION LAB

THUNDERBOOM

MUSIC TECH

INNOVATION LAB

THUNDERBOOM

MUSIC TECH

INNOVATION LAB

THUNDERBOOM

MUSIC TECH

INNOVATION LAB

THUNDERBOOM

MUSIC TECH

INNOVATION LAB

THUNDERBOOM

MUSIC TECH

INNOVATION LAB

SAAZ-E-YAAD

SAAZ-E-YAAD

SAAZ-E-YAAD

Saz-e-yaad is a standalone desktop application for generating music inspired by Kashmiri folk
traditions. It combines several AI models into a single interface, allowing users to create new
musical ideas from scratch, from an existing audio clip, or from a text prompt.


Unlike a traditional synthesizer or sampler, Saz-e-yaad generates entirely new audio using
machine learning models trained on curated recordings of Kashmiri folk music.

 

 

Installation


The current public release is available for macOS on Apple Silicon.
To install the application:

  1.  Download the latest Saz-e-yaad release.
  2. Open the application on an Apple Silicon Mac.
  3. Launch the desktop application.

The available documentation does not describe additional installation or configuration steps
beyond downloading and launching the application. The project source code is available on
GitHub for users who wish to explore or build the project themselves.

 


Choosing a model


The application includes several specialised AI models, each trained on a different version of
the dataset.


Full – Generates complete musical textures, including melody, harmony and
accompaniment.
Harmony – Focuses on melodic and harmonic material without percussion.
Rhythm – Trained on isolated rhythmic material.
Vocals – Trained on vocal recordings from the curated dataset.

Each model produces a different type of musical output depending on the material it learned
during training.

 


Generating audio


Once a model has been selected, there are three ways to start a generation:

Unconditional – Generate a completely new piece without any input.
Audio prompt – Drag an audio file into the application to guide the generation.
Text prompt – Describe the desired result with text and let the model generate audio
based on that prompt.


While the model is generating, decoded audio is streamed into a live playback window. Instead
of waiting for the final render, you can hear the composition gradually take shape.

 


Adjusting the generation


Several controls allow you to influence the generated result.


Available parameters include:
● Output length
● Diffusion steps
● Temperature
● Surprise (seed scale)
● Number of generated samples
● Optional random seed for repeatable results

Completed generations appear in the sample panel and can be exported as standard .wav files
for use in any digital audio workstation (DAW).

 


Continuing a generation


One of the application's most useful features is Continue Selected.


Instead of starting over, you can select a previously generated sample and continue generating
from that point. The selected audio is used as the prefix for the next generation, making it
possible to gradually develop longer musical ideas through multiple iterations.

 


Example workflow


A producer wants to create an atmospheric introduction for a new track.

  1.  Select the Harmony model.
  2. Choose Unconditional generation.
  3. Generate several short musical ideas.
  4. Listen to the generated samples and select the most interesting one.
  5. Use Continue Selected to expand that idea into a longer composition.
  6. Export the final result as a WAV file.
  7. Import the file into Ableton Live (or another DAW) for further editing, arrangement or processing.

 

This workflow reflects the intended use of the application: generating original musical material
that can serve as the starting point for music production rather than producing a finished song
automatically.

 


Under the hood


Saz-e-yaad combines several modern machine learning components inside a single desktop
application.


The application performs inference using TorchScript, encodes and decodes audio with
music2latent, uses CLAP for audio and text conditioning, and runs on Apple Silicon through
Metal Performance Shaders (MPS), automatically falling back to the CPU when necessary. The
AI models themselves are based on Continuous Autoregressive Models (CAM) operating in
latent audio space with diffusion-based sampling.

 

Related Content

Open Culture Tech

Open Culture Tech

Open Culture Tech makes new technology, such as AI and holograms, accessible to artists in The Netherlands by developing and sharing publicly available tools, showcases and knowledge.