Repository navigation
Attribute(s) for the size tag
#260
Replies: 3 comments
|
Thanks for this suggestion, @gramian! We're also thinking about how we can make it easier to parse the contents of the There are a couple adjacent discussions you may be interested in—if you have any feedback on these ideas, we'd love to hear from you:
|
|
Cross posting my comment from #242 (comment)
|
|
Could the proposal include lossless interpretation fixtures, alongside the possible unit attributes? They would make the byte-extraction use case testable without restricting Size to numeric content. The current 4.7 Size definition is optional, repeatable free text and includes physical extent and duration. NIST's unit definitions distinguish decimal MB from binary MiB. However, interpreting a symbol according to those definitions does not prove that an unknown legacy depositor intended that convention, or that the original measurement was exact. For a declared SI/IEC interpretation, a conservative byte-unit-only parser could use these fixtures:
Here is a narrow synthetic downstream example, not a DataCite validator or proposed XML representation. The caller must establish the convention; it is never inferred merely from the string. If that convention is unresolved, even import re
from fractions import Fraction
def interpret_size(raw, unit_profile=None):
result = {"raw": raw, "unit_profile": unit_profile, "bytes": None}
if unit_profile != "SI/IEC":
return {**result, "reason": "unit convention not established"}
match = re.fullmatch(
r"([0-9]+(?:\.[0-9]+)?)[ \t]+(B|kB|MB|GB|TB|KiB|MiB|GiB|TiB)",
raw.strip(),
)
if not match:
return {**result, "reason": "outside declared byte-unit grammar"}
factors = {"B": 1, "kB": 10**3, "MB": 10**6, "GB": 10**9, "TB": 10**12,
"KiB": 2**10, "MiB": 2**20, "GiB": 2**30, "TiB": 2**40}
value = Fraction(match[1]) * factors[match[2]]
if value.denominator != 1:
return {**result, "reason": "non-integer byte result; do not round"}
return {**result, "bytes": int(value), "reason": "derived under declared units"}I checked the shown cases plus unit, whitespace and precision boundaries locally: 26 synthetic inputs, preserving the raw value and declining conversion when the convention is unestablished. This does not validate a deposit, cover all legacy notation or measure a file's actual payload. Would the proposal preserve the original extent and make the conversion convention and unresolved state explicit? An unsupported value in this example is not a DataCite validation error. The Distribution discussion #211 remains a proposal under revision; these fixtures do not assume adoption of Disclosure: I am founder/chairman of Bharath Healthcare (BHCL). This source-based methods contribution was prepared with Codex assistance; it does not seek endorsement of our work. |
Uh oh!
There was an error while loading. Please reload this page.
What is the problem that your suggestion solves?
Currently, the
sizefield is hard to parse free-text. Specifically for my case, I try to find the size in Bytes (if available) of DataCite described resources, this means I need to filter for all variants to write a digital information unit. So in addition to check for all possible writings of decimal and binary prefixes for Bytes, I need to decipher if the full unit name is used (in some variants).What solution might meet your needs?
I would most prefer to have the
sizefield content to be numeric and have one or more attributes describing its content, for example justunitas a sole (free-text) attribute would be already helpful, an additional (optional) attribute for digital content likebyteswould be the best case for me, where I would just have to parse the prefixes.Your name
No response
Your organization
No response
What alternatives have you tried or considered?
No response
Is there anything else you would like to share?
No response
What group(s) would benefit from your suggestion?
If other group(s), please describe.
No response
All reactions