Repository navigation
Request for comment: New Distribution property #211
Replies: 12 comments 10 replies
|
This nicely maps to DCAT class Distribution with properties
I'm not sure about |
|
@KellyStathis Thanks for starting this discussion! Could you please elaborate on how Distribution would relate to the resourceType for dataset files I suggested some time ago (see Dataverse GitHub issue #5086 and PID Forum thread)? |
|
Thanks for clarifying, @KellyStathis! I think Dataverse already provides the relationType information. However, in DataCite/DataCite Commons dataset and file-level DOIs would all be listed at the same level, which is somewhat confusing, because it might make people think our repository has more than 100.000 dataset, but actually most of these DOIs are at file-level. Will the Distribution property enable distinguishing between datasets and files? |
|
Thanks, @KellyStathis! I think introducing a new resourceTypeGeneral for "File" or "Component" or similar could do the job. That's basically, what I suggested back in 2018. What would be the right way to proceed with this suggestion? |
|
One thing that's come up for us in considering whether/how we might include metadata in this new property for our own DOIs, is the impact of their being MANY files. We have objects with almost 50,000 files. That's likely to cause an issue with payload size for us or you at some point, it seems like. Or just in the time it takes to generate the metadata or possibly even worse -- checking and updating it whenever a work is modified to see if, for example, a file size has changed, or a single file has been added. I don't have any suggestions or request, just suggesting that this point be considered. Thanks! |
|
I would like to see some thought given to how to support file directories on the cloud (e.g., s3), particularly when some additional information is needed to access the file (e.g., access keys / similar). Any ideas on that? |
|
There has always been this strange inference in the DataCite metadata schema that the dataset being identified consists of a single file, or else several files that can be treated as a single file because they all have the same format/access conditions/usage rights. That does not reflect the datasets in our repository. We warmly welcome this proposal, which would avoid the need for the ugly workarounds we are currently using: "relatedIdentifiers": [
{
"relatedIdentifier": "https://url/for/filename.ext",
"relatedIdentifierType": "URL",
"relationType": "HasPart"
}
],
"formats": [
"filename.ext = mime/type"
],
"sizes": [
"filename.ext = 196kB"
],However, excluding rights is a missed opportunity, because we assign licences per-file not per-dataset, so even with this suggestion we would still be left with an ugly workaround: "rightsList": [
{
"rights": "filename.ext is licensed under the XXX License",
...
}
], |
|
Thanks everyone for your feedback and contributions to this Request for Comment. I’m writing with an update on the status of Distribution and a request for additional feedback. We have an updated draft of the Distribution property available here: https://datacite-metadata-schema.readthedocs.io/en/4.8_draft/properties/distribution/ as part of the working draft of Schema 4.8. While preparing examples to accompany the Distribution property, several important issues have surfaced in discussions with the Metadata Working Group. First, the current design for Distribution centers around a stable direct download link, with contentURL as the main required sub-property. However, we’ve found that many repositories don’t have stable direct download links. A common pattern, for example, is to use a dynamically generated, expiring download link. This approach is implemented by Dataverse software (here’s an example: https://doi.org/10.5683/SP3/KPAQRO) among others. Expiring URLs like these wouldn’t be actionable contentURLs. We’re also concerned that stable contentURLs might become rarer as repositories deliberately seek to prevent machine access. Many are concerned about network traffic, and with increasing strain on collections due to AI harvesting1, the landscape has changed considerably since we first started working on Distribution four years ago. We have a few questions to help move this forward—thanks in advance for your input! We want to make sure that Distribution is a useful addition to the DataCite Metadata Schema.
Footnotes
|
|
In my opinion, the distribution property is unneeded if reporistories would better implement FAIR signposting. Basically at PANGAEA you can resolve the DOI with a HEAD request (following redirects) of course and all alternatie ditributions are announced as HTTP Link headers as decribed in Content Negotiation RFCs and FAIR Signposting. So I am against this change. Adding another layer of distribution URLs in DataCite metadata makes it even more complicated for a consumer to find the correct link to download data or alternative metadata representations. FAIR signposting was made exactly for that use case, so a user should only need to resolve the DOI and query the repository for alternative links with standard HTTP requests. If you download the landing page, the links are also reported as HTML So I would not like to move that to metadata. Especially as links to data or other alternative representations are often dynamic and not stable so should not be added to metadata. |
|
I will reiterate that the thing I welcome most about this development is the ability to associate file-specific metadata (filename, byte size, media type, access level, and potentially others) with specific files. It just makes so much more sense than specifying file-specific metadata at the whole-dataset level. The problem with the proposed implementation is that all this file-specific goodness is provided as a happy by-product of providing stable direct download URLs: the ContentURL is made primary and everything else has to hang off it. If a repository cannot or does not want to publish direct download links, they are then unable to provide any other file-specific metadata even though it would still make sense for them to do so. To be clear, I don't think it's a valid argument that, because some repositories can't or don't want to provide stable direct download URLs, no repository should be able to, nor that because each repository could potentially implement their own technologies for providing machine-readable direct download URLs, that there's no benefit to having a simple standard way of providing them that works for everyone and doesn't involve having to visit each repository individually (and craft a specific approach). So to answer your questions:
In short, the structure "resource – component – download info" makes more logical and pragmatic sense than "resource – download info – other component info", so flipping the emphasis around would be a good move. |
|
An additional optional property would be a typed checksum for the file content e.g: {
"contentUrl": "https://example.org/readme.txt",
"byteSize": 838861,
"mediaType": "text/plain",
"accessType": "open",
"contentName": "readme.txt",
"checksum": "801a9be154c78caa032a37b4a4f0747f1e1addb397b64fa8581d749d704c12ea",
"checksumType": "SHA256",
"checksumTypeUri": "http://id.loc.gov/vocabulary/preservation/cryptographicHashFunctions/sha256"
} |
|
Hi everyone, thanks for sharing your feedback on this proposal. We’ve heard from many in the community that the proposed implementation would make it easier both to share content URLs and describe individual files. At the same time, this round of feedback revealed challenges with how the design anchors file-specific metadata to a required content URL. This is especially true for the XML structure, in which byteSize, mediaType, etc. were attributes of the contentURL element. This doesn’t provide flexibility to expand or adjust the structure later—for example, we wouldn’t be able to make contentURLs optional, or add sub-elements with their own attributes. Because we keep all minor version changes backward-compatible, it’s important that when we implement a new property, we start with a design that can grow to meet community use cases as they evolve. We’re going to take this back to the DataCite Metadata WG to revise, and will share another proposal in this space for discussion. Thanks for working with us on this—more to come. |
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Currently, machines have no standardized way to retrieve resources identified with DataCite DOIs. When a DOI resolves to a landing page, machines need to rely on machine-readable file download information being embedded in that landing page, or other means that are inconsistent across repositories.
To address this challenge, DataCite is proposing a new Distribution property for the DataCite Metadata Schema.
Proposed new Distribution property
21. Distribution
21.1 contentURL
https://example.org/data.csvftp://example.org/data.txthttps://example.org/files.gzip21.1.a byteSize
1048576(for 1 Megabyte)21.1.b mediaType
application/zipaudio/mpeg21.1.c accessType
OpenEmbargoedRestricted21.1.d contentName
Definition: A name given to the content at the contentURL.
readme.txtExamples
XML
JSON
Providing feedback
We are interested in your thoughts on this proposal and possibilities for improvement. Contribute your feedback by leaving a comment on this post. You are welcome to reply by answering any or all of the questions below:
If you would like to share feedback without creating a GitHub account, you can email us at support@datacite.org and indicate whether you’d like us to post a comment on your behalf.
This discussion will close on 16 January 2026 (extended from 19 December 2025) , but you can still submit feedback on DataCite Suggestions.
All reactions