1Introduction
1.1Music Transparency and Provenance Challenge
This guidance explains how Content Credentials can be applied to music assets to support transparent provenance across creation, production, delivery, and distribution workflows.
The guidance focuses on three core music use case families:
- AI disclosure and labeling
- Enabling verifiable statements about whether and how AI was used in the creation or production of music. This helps ensure information about AI use is available to downstream stakeholders, including platforms, rightsholders, creators, and consumers, to inform AI disclosure and content labeling decisions.
- Copyrightability and Chart Eligibility
- Documenting how music was created, including whether it was created by human creators, generated by AI systems, or produced through a combination of both. This information may support downstream activities involved in evaluating copyrightability and chart eligibility.
- Credits
- Maintaining information about creator identities, ingredient origins, and production history throughout the lifecycle of music. This helps ensure creators are properly credited for their contributions through a verifiable provenance record of the actions applied to the content.
These use cases reflect common scenarios where provenance information will help describe how a music asset was created, modified, reviewed, or prepared for release. For additional guidance on applying AI disclosure with Content Credentials, see the C2PA AI Labelling Guidance.
Content Credentials and the associated manifest do not make judgments about artistic merit, copyright status, ownership, chart eligibility, licensing, or the extent of AI use. Rather, they provide a clear, verifiable provenance record of how the music was created and modified.
1.2Content Credentials
Content Credentials provide a way to associate signed provenance information with a digital asset. In the C2PA model, a claim generator creates a C2PA Manifest containing assertions about the asset, such as actions, ingredients, software information, timestamps, and signatures. In a music workflow, a claim generator could be a DAW (Digital Audio Workstation), plug-in, AI music tool, mastering system, distributor service, or other software component that creates or updates provenance information as the asset moves through production and delivery.
A validator is a tool or service that reads the manifest and checks whether the signed information is well formed, has not been tampered with, and was signed using a trusted credential. For example, a label might use a verifier to review a submitted audio asset before accepting it into a release workflow.
This means Content Credentials are primarily an interoperability layer between tools and services. Provenance data can be generated, signed, preserved, read, and validated by systems across the music workflow. Data interoperability between C2PA and specifications like CAWG and DDEX allows publishers and platforms to incorporate this information into rights clearing, payments, and other critical operational systems. Selected information may be surfaced by user-facing services according to their own display, policy, privacy, and user-experience choices.
Content Credentials rely on a binding between the C2PA Manifest and an asset. For audio assets, this binding can be implemented through direct embedding or external manifest reference (sidecar). When direct embedding is used, implementers should use the container mechanism appropriate to the audio format. For example, ID3v2-compatible compressed audio files, such as MP3, can use an ID3 General Encapsulated Object (GEOB) frame to carry the C2PA manifest payload. For additional format-specific details, implementers should refer to the C2PA Specification embedding annex.
The claim generator signs the manifest using a signing credential, typically an X.509 certificate associated with the organization, service, or software component operating the claim generator. A validator can then check the manifest structure, the claim signature, the asset binding, and whether the signing certificate chains to a trust anchor on the trust list.
This trust model does not prove that every assertion is complete, accurate, licensed, or legally sufficient. It provides tamper-evident provenance that downstream systems can verify and interpret according to their own trust, policy, display, and business rules.
1.3Who This Guidance Is For
This guidance is intended for participants across the music ecosystem who create, process, distribute, receive, display, or verify music assets and related provenance information.
This includes DAWs, plug-in developers, instrument makers, AI music tools, mastering services, labels, distributors, DSPs, metadata providers, and other services involved in the lifecycle of a music asset.
The goal is to support interoperability across this chain. Content Credentials can help carry provenance information from creation and production through delivery and distribution, while allowing implementations to decide what information is recorded, preserved, displayed, or withheld according to policy, privacy, business, and user-experience needs.
This guidance should be understood as a first-phase approach focused on manifesting music assets at the track, stem, and major-component level. It does not attempt to describe every subcomponent within a stem, such as individual notes, automation data, or micro-level production decisions. Future guidance may address more granular provenance patterns as music workflows, tooling, and implementation practices mature.
Where an implementer wants to assert the identity of a person, company, label, distributor, creator, or other actor, this guidance recommends using CAWG identity assertions: CAWG Specifications.
This guidance is intended to complement existing music-industry metadata and delivery standards, including DDEX-based workflows. It does not replace rights databases, contracts, licensing records, copyright analysis, royalty systems, or platform policy decisions.
2Core Concepts for Music Assets
2.1Actions and digitalSourceType
C2PA actions describe events or operations that happened to an asset, such as creation, editing, enhancement, mixing, mastering, transcoding, etc.
The actions c2pa.created and c2pa.opened are fundamental to understanding how the C2PA specification applies to audio files. c2pa.created indicates that an asset was created, while c2pa.opened indicates that an existing asset was opened for further processing. These two actions are not interchangeable. For example, if a user opens an existing track in a DAW and then exports it as a new file, the export action should be recorded as c2pa.created, while the opening of the original track should be recorded as c2pa.opened.
In the context of C2PA, the term “created” simply means an asset came into existence, regardless of whether it was created by a human, AI-generated, or created with AI assistance. While creation is traditionally a human endeavor, the standard uses the term to also describe computer-generated files and AI-assisted workflows.
The digitalSourceType of an action represents a controlled vocabulary used to describe how that action was performed. For example, it can indicate whether the action involved digital capture, human editing, non-AI algorithmic enhancement, trained algorithmic media (AI generated), or a composite process.
In a C2PA actions assertion, digitalSourceType is recorded on an individual action. It should therefore be interpreted together with the action it is attached to, not as a standalone classification for the entire asset.
The broader production history should be expressed through the sequence of actions, ingredients, software or tool information, timestamps, and other relevant assertions.
| C2PA Action | Possible Digital Source Type | Music / Audio Examples |
|---|---|---|
c2pa.created | digitalCapture, digitalCreation, trainedAlgorithmicMedia, or computationalCapture | Recording a vocal, creating a synth part, generating an AI stem, or computationally capturing audio. |
c2pa.opened | Usually determined by the parent ingredient and subsequent action | Opening an existing track, stem, master, or recording as the starting point for a new manifest. |
c2pa.edited | humanEdits or compositeWithTrainedAlgorithmicMedia | Human editing, comping, timing edits, or edits involving trained algorithmic media. |
c2pa.placed | digitalCapture, compositeSynthetic, or trainedAlgorithmicMedia | Placing a sample or loop into a mix. |
c2pa.enhanced | algorithmicallyEnhanced or compositeWithTrainedAlgorithmicMedia | Non-AI cleanup, restoration, leveling, or AI-based voice cleaning and enhancement. |
c2pa.mixed | humanEdits or compositeSynthetic | Mixing stems manually, or combining human/captured material with AI-generated material. |
c2pa.mastered | humanEdits or algorithmicallyEnhanced | Human mastering or non-AI algorithmic mastering. |
c2pa.transcoded | Usually no new digitalSourceType | Format conversion, such as WAV to MP3, where the audio content is not editorially changed. |
When a c2pa.created action is used to identify media created using a generative AI model, a digitalSourceType of trainedAlgorithmicMedia is appropriate. Other types of actions that involve algorithmic or machine-learning-based enhancement that do not generatively alter the main content may instead be described as algorithmicallyEnhanced. When an action combines multiple discrete elements, such as c2pa.placed, and you need to indicate that at least one element was created using trained algorithmic media, compositeSynthetic may be appropriate, though the use of the digitalSourceType field on the placed ingredient is also useful in this case.
When an action uses a generative AI model to augment, correct, or enhance existing media, such as inpainting, outpainting, or replacing a segment of a recording, compositeWithTrainedAlgorithmicMedia is appropriate. When an action uses non-AI algorithmic processing, algorithmicallyEnhanced shall be appropriate and should not be used for generative AI processing. For further detail, see the AI labelling guidance document.
Example: AI-Based Voice Cleaning
The following simplified example shows how an actions assertion could describe an existing vocal recording that was enhanced using an AI-based voice cleaning tool.
{
"assertions": {
"c2pa.actions.v2": {
"actions": [
{
"action": "c2pa.opened",
"when": "2026-04-23T09:14:00-07:00",
"description": "Opened existing vocal recording for editing."
},
{
"action": "c2pa.enhanced",
"when": "2026-04-23T09:15:00-07:00",
"description": "Applied AI-based voice cleaning and noise reduction.",
"digitalSourceType": "http://cv.iptc.org/newscodes/digitalsourcetype/algorithmicallyEnhanced ",
"softwareAgent": "AI Voice Cleaning Tool"
}
]
}
}
}In this example, digitalSourceType is attached to the c2pa.enhanced action and describes how that enhancement action was performed. The original vocal recording should be represented as an ingredient where possible.
Example: AI-Based Voice Cloning
The following simplified example shows how an actions assertion could describe a vocal created using an AI-based voice cloning tool.
{
"assertions": {
"c2pa.actions.v2": {
"actions": [
{
"action": "c2pa.created",
"when": "2026-04-23T09:15:00-07:00",
"description": "Applied AI-based voice cloning.",
"digitalSourceType": "http://cv.iptc.org/newscodes/digitalsourcetype/trainedAlgorithmicMedia",
"softwareAgent": "AI Voice Cloning Tool"
}
]
}
}
}In this example, digitalSourceType is attached to the c2pa.created action and describes how that created action was performed. The original vocal recording should be represented as an ingredient where possible.
2.2Ingredients
An ingredient is a source asset used to create or modify another asset. In music, this may include a vocal recording, stem, take, sample, loop, AI-generated layer, previous mix, master, or audio extracted from video.
Ingredients are useful when a final music asset combines multiple sources. For example, if a human vocal is mixed with an AI-generated instrumental, both should be represented as ingredients.
Actions describe what happened, digitalSourceType describes how a specific action was performed, and ingredients describe what source assets were used.
Example: Final Mix with Stem Ingredients
The following simplified example shows how a final mix could reference multiple stem ingredients.
{
"assertions": {
"c2pa.ingredient.v3": {
"data": {
"alg": "sha256",
"hash": "76d9aff0167492306864be6d30cb16f71e4c0be650f408889c3d5c90e7f39acb",
"url": "https://fabrikam.com/session/LeadVocal_Stem.wav"
},
"dc:format": "audio/wav",
"dc:title": "LeadVocal_Stem",
"instanceID": "urn:c2pa:4ec28767-3353-4a62-80ce-c56548e3ce9c",
"relationship": "componentOf"
},
"c2pa.ingredient.v3__1": {
"data": {
"alg": "sha256",
"hash": "f94de83ad6a029a13a71ad12fdebea16fe7bc089814588949e5e7f75091b5e06",
"url": "https://fabrikam.com/session/Drums_Stem.wav"
},
"dc:format": "audio/wav",
"dc:title": "Drums_Stem",
"instanceID": "urn:c2pa:52938c01-ab45-4a3a-97f0-71cbc60f6b98",
"relationship": "componentOf"
},
"c2pa.ingredient.v3__2": {
"data": {
"alg": "sha256",
"hash": "4ece6641679d841d2f0c702e93c6c8285e43eebfd0d47958e0feee0f7f674847",
"url": "https://fabrikam.com/session/AI_Pad_Stem.wav"
},
"dc:format": "audio/wav",
"dc:title": "AI_Pad_Stem",
"instanceID": "urn:c2pa:25a70ca9-66a8-4a42-8537-ee8dee7a6d16",
"relationship": "componentOf"
}
}
}In this example, the final mix has three ingredients: a lead vocal stem, a drum stem, and an AI-generated pad stem. Each ingredient can be referenced, hashed, and described separately.
2.3AI Disclosure Assertion
The 2.4 revision of the C2PA Specification introduces the c2pa.ai-disclosure assertion. It provides for a structured, machine-readable way to disclose information about AI involvement in an asset, such as model identification and human-oversight or review information.
c2pa.ai-disclosure is complementary to, not a replacement for, actions, digitalSourceType, and ingredients. Each serves a distinct purpose:
- digitalSourceType: describes the nature of a given action (when used in an actions assertion) or a specific ingredient (when used in an ingredient assertion). It can be used to indicate whether it was captured, human-edited, or generated using trained algorithmic media.
- Ingredients: describe the source or additional assets and inputs used to produce an asset.
- c2pa.ai-disclosure: provides structured information about the AI model and generation process, such as model identification and human-oversight information.
This guidance does not require implementers to include a c2pa.ai-disclosure assertion. Where richer, structured disclosure of the AI model or generation process is useful, for example to support AI disclosure and labeling use cases, implementers should consult the C2PA Specification for the assertion’s full field definitions and requirements.
2.4Region of Interest
A C2PA assertion may apply to only part of an asset. For music and audio assets, this can be useful when an action affects only a specific time range, such as an AI-replaced vocal phrase, localized noise reduction, stem cleanup, or generative extension. Further details can be found in the Region of Interest section of the specification.
A temporal region can identify the affected segment of the recording.
Example: Audio Temporal Region for AI-Replaced Vocal Phrase
The following simplified example shows how a temporal region could be used to identify the part of a recording affected by an AI-assisted edit.
{
"actions": [
{
"action": "c2pa.edited",
"when": "2025-04-01T10:01:30Z",
"softwareAgent": { "name": "ExampleGenAI Studio", "version": "3.1.0" },
"digitalSourceType":
"http://cv.iptc.org/newscodes/digitalsourcetype/compositeWithTrainedAlgorithmicMedia",
"changes": [
{
"region": [
{
"type": "temporal",
"time": {
"type": "npt",
"start": "68.00",
"end": "72.00"
}
}
],
"name": "AI-Replaced Vocal Phrase",
"identifier": "modified-audio-region-001",
"description":
"Vocal phrase replaced by AI during seconds 68.00 to 72.00 of the recording."
}
]
}
]
}This allows implementers to describe localized provenance without implying that the same action applies to the entire asset.
2.5Manifests
A manifest is the cryptographically signed wrapper that contains a list of ingredients and the associated operations performed on them in the form of actions. A hypothetical C2PA compliant vocal recording tool would insert a manifest into the vocal audio at the time of creation. If that vocal is modified by a second compliant tool (say an effects processor), that tool generates a new manifest that securely references the first manifest, adding a record of the effects processing. The manifests contain cryptographic hashes, allowing subsequent modifications to the referenced audio or metadata to be detected. If a compliant DAW combines the processed audio with an instrumental performance and an AI generated stem (also created by C2PA compliant tools), the DAW will create a manifest that securely references the manifests of all three ingredients.
In this way, sequential changes and the combining of assets can all be securely tracked.
3Example Workflow
3.1Scenario 1: AI-Generated Instrumental Combined with Human Vocal
A producer creates an instrumental track using an AI music generation tool. The producer then records a human vocal performance directly in a DAW.
Inside the DAW, the producer combines the AI-generated instrumental and the human vocal recording, edits and mixes the sources into a final track, then exports the finished audio file for delivery to a label, distributor, or DSP.
In this workflow, the AI-generated instrumental and the human vocal should each be represented in their own manifest and referenced as ingredients in the final mix’s manifest. An actions assertion cannot contain more than one c2pa.created or c2pa.opened action, since that action represents how the asset described by that manifest came into existence. The creation of the instrumental and the vocal should therefore be recorded in their own manifests rather than as multiple creation actions inside the final mix’s manifest. However, the manifests of both will be wrapped into the manifest of the final mix.
Manifest Excerpt: AI Instrumental
The following simplified excerpt shows the manifest for the AI-generated instrumental.
{
"assertions": {
"c2pa.actions.v2": {
"actions": [
{
"action": "c2pa.created",
"description": "AI-generated instrumental.",
"digitalSourceType":
"http://cv.iptc.org/newscodes/digitalsourcetype/trainedAlgorithmicMedia",
"softwareAgent": {
"name": "AI Music Tool Fabrikam",
"version": "1.0"
}
}
]
}
},
"claim.v2": {
"claim_generator_info": {
"name": "AI Music Tool Fabrikam",
"version": "1.0",
"specVersion": "2.4.0"
}
}
}Manifest Excerpt: Human Vocal
The following simplified excerpt shows the manifest for the recorded human vocal performance.
{
"assertions": {
"c2pa.actions.v2": {
"actions": [
{
"action": "c2pa.created",
"description": "Recorded human vocal performance.",
"digitalSourceType":
"http://cv.iptc.org/newscodes/digitalsourcetype/digitalCapture",
"softwareAgent": {
"name": "DAW",
"version": "1.0"
}
}
]
}
},
"claim.v2": {
"claim_generator_info": {
"name": "DAW",
"version": "1.0",
"specVersion": "2.4.0"
}
}
}Manifest Excerpt: Final Mix
The following simplified excerpt shows the manifest for the final exported mix. It references the AI instrumental and the human vocal as ingredients, rather than repeating their creation as separate c2pa.created actions.
In this example, softwareAgent identifies the software used to perform a specific action, while claim_generator_info identifies the software that generated the claim or manifest.
{
"assertions": {
"c2pa.actions.v2": {
"actions": [
{
"action": "c2pa.created",
"description": "Exported final mixed track.",
"digitalSourceType": "http://cv.iptc.org/newscodes/digitalsourcetype/compositeSynthetic",
"softwareAgent": {
"name": "DAW",
"version": "1.0"
}
},
{
"action": "c2pa.placed",
"description": "Placed AI instrumental and human vocal into the mix.",
"digitalSourceType": "http://cv.iptc.org/newscodes/digitalsourcetype/humanEdits",
"parameters": {
"ingredients": [
{
"url": "self#jumbf=c2pa.assertions/c2pa.ingredient.v3",
"alg": "sha256",
"hash": "<ai-instrumental-assertion-hash>"
},
{
"url": "self#jumbf=c2pa.assertions/c2pa.ingredient.v3__1",
"alg": "sha256",
"hash": "<vocal-ingredient-assertion-hash>"
}
]
}
},
{
"action": "c2pa.mixed",
"description": "Mixed AI instrumental with human vocal.",
"digitalSourceType": "http://cv.iptc.org/newscodes/digitalsourcetype/humanEdits"
}
]
},
"c2pa.ingredient.v3": {
"dc:title": "AI Instrumental",
"dc:format": "audio/wav",
"relationship": "componentOf",
"activeManifest": {
"url": "self#jumbf=/c2pa/<ai-instrumental-manifest>",
"alg": "sha256",
"hash": "<ai-instrumental-manifest-hash>"
}
},
"c2pa.ingredient.v3__1": {
"dc:title": "Human Vocal",
"dc:format": "audio/wav",
"relationship": "componentOf",
"activeManifest": {
"url": "self#jumbf=/c2pa/<vocal-manifest>",
"alg": "sha256",
"hash": "<vocal-manifest-hash>"
}
}
},
"claim.v2": {
"claim_generator_info": {
"name": "DAW",
"version": "1.0",
"specVersion": "2.4.0"
}
}
}Interpretation
This example shows how the exported file can contain a new active manifest for the final mix while preserving the AI instrumental and human vocal as separate ingredient manifests linked through hashed references.
The AI instrumental’s manifest contains a single c2pa.created action with digitalSourceType trainedAlgorithmicMedia, indicating that the instrumental was created using trained algorithmic media.
The human vocal’s manifest contains a single c2pa.created action with digitalSourceType digitalCapture, indicating that the vocal was recorded from a real-world source.
The c2pa.mixed action describes the later combination of the AI-generated instrumental and the human vocal. Its digitalSourceType is humanEdits, indicating that the mixing itself was performed manually by a human in the DAW, even though the mixed sources include trained algorithmic media. compositeWithTrainedAlgorithmicMedia is reserved for actions that use generative AI to augment or edit existing media, such as inpainting, outpainting, or replacing a segment of a recording, and is not the correct value for describing a manual mix of separately created sources.
This example is intentionally simplified. A full implementation may also include ingredient hashes and URLs, signatures, validation results, identity assertions, soft binding, or links to related metadata systems.
3.2Scenario 2: Licensed Remix in a Platform Environment
A user accesses a licensed song inside a platform that is authorized to make a remix. The user creates a remix using the platform’s built-in tools. Note that the portability of the resulting remix may be limited by the platform. “Walled-garden” systems may not allow for the downloading and distribution of the remix.
In this workflow, the original licensed recording is the previous asset from which the remix is derived, so it should be represented as a parentOf ingredient and referenced from the c2pa.opened action through parameters.ingredients. The remix action can be described using c2pa.remixed, with tool, timestamp, and platform information recorded as appropriate.
Content Credentials can help describe that a remix action occurred and identify the source asset and platform context, but they do not replace the underlying license, platform terms, or rights-management system.
Manifest Excerpt
The following simplified excerpt shows how a licensed AI remix workflow may record that a remix action occurred inside a platform environment.
{
"assertions": {
"c2pa.actions.v2": {
"actions": [
{
"action": "c2pa.opened",
"description": "Opened licensed source recording inside the platform.",
"parameters": {
"ingredients": [
{
"url": "self#jumbf=c2pa.assertions/c2pa.ingredient.v3"
}
]
},
"softwareAgent": {
"name": "Remix Platform",
"version": "1.0"
}
},
{
"action": "c2pa.remixed",
"description": "Created an AI-generated remix using platform-provided tools.",
"digitalSourceType":
"http://cv.iptc.org/newscodes/digitalsourcetype/trainedAlgorithmicMedia",
"softwareAgent": {
"name": "Remix Platform",
"version": "1.0"
}
}
]
},
"c2pa.ingredient.v3": {
"dc:title": "Song title A - Artist X",
"dc:format": "audio/wav",
"relationship": "parentOf"
}
},
"claim.v2": {
"claim_generator_info": {
"name": "Remix Platform",
"version": "1.0",
"specVersion": "2.4.0"
}
}
}Interpretation
This example shows how Content Credentials can describe an AI-assisted remix of a licensed source recording within a platform workflow.
The c2pa.opened action describes the licensed source recording being opened inside the remix platform. Because the remix is derived from this recording, the source recording is represented as a parentOf ingredient and referenced directly from the c2pa.opened action through parameters.ingredients.
The c2pa.remixed action describes the creation of the AI-generated remix. Its digitalSourceType is trainedAlgorithmicMedia, indicating that the remix action involved both the licensed source recording and trained algorithmic media.
This example does not attempt to express the license terms, platform permissions, royalty treatment, or downstream usage rules. Those remain handled by the platform’s licensing, rights-management, policy, and commercial systems outside the C2PA manifest.
Note that the manifest will be the same regardless of whether the platform is open or a “walled garden” system.
4Annexes
4.1DDEX and Metadata Alignment
Content Credentials act as an upstream provenance layer that supports DDEX metadata. DDEX remains the exchange format for structured music metadata between labels, distributors, DSPs, and other commercial partners. A claim generator should capture verifiable information during creation and production, such as actions, ingredients, tools, timestamps, and digitalSourceType values. This information may later be used by record labels to populate, validate, or support DDEX metadata.
4.2Persistence and Recovery with Soft Binding
Content Credentials may be stripped during upload, resizing, transcoding, etc. C2PA soft binding can support recovery by resolving a decoupled manifest through a manifest repository, provided the asset was prepared with a suitable soft binding signal and the manifest was registered. The C2PA soft binding specification describes this model as a way to recover manifests that have become decoupled from their assets. The C2PA Specification also defines optional soft binding approaches, including watermarking and fingerprinting, that can support indirect lookup of a Content Credential when direct embedding is not available or when embedded metadata may not survive downstream processing.
For music assets, soft binding may be useful because audio often passes through transformations such as transcoding, loudness normalization, clipping, platform ingestion, user-generated content workflows, and redistribution. Soft binding is optional and this guidance does not require any music company, platform, distributor, or tool provider to implement watermarking or fingerprinting.
4.3Best Practices for Audio Watermarking
Where watermarking is used with Content Credentials applied to music, implementers should prefer technologies that are imperceptible to listeners and robust under real-world distribution conditions.
Relevant considerations include:
- Imperceptibility to listeners
- Robustness after transcoding, resampling, clipping, loudness normalization, and format conversion
- Detection from short excerpts where appropriate
- Low false positive and false negative rates
- Reliable detection after common streaming and social-platform processing
- Clear behavior when multiple watermarked assets are mixed, overlaid, or re-watermarked
- Support for lookup-based payloads, where the watermark carries a compact identifier rather than full metadata
For catalog-level identification, some implementations may choose to use an identification payload, such as a 32-bit lookup key, to connect the detected watermark to a manifest repository. The appropriate payload size is an implementation and business decision, not a C2PA requirement.
Companies developing watermarking or fingerprinting technologies for Content Credentials should consider participating in the C2PA soft binding ecosystem and requesting inclusion in the relevant soft binding algorithm list where appropriate.