Repository navigation
Provide the PG Schema #189
Description
Activity
Hi, thanks for sharing that. I had no idea about PG-Schema.
Indeed this could help our users, but I want to point out that IYP schema is quickly evolving (adding new datasets or datasets changing over time) so automation is probably needed. But I already foresee problems with automation, we import a lot of properties from the different datasets we ingest. Some of these properties are very important and our code make sure that they are of the correct type when we import them, and some properties are just 'copy/paste' from the datasets we ingest. I don't think the later ones should appear in the schema.
Currently we are documenting important properties in our documentation for nodes and relationships.
You're right — evolving schemas and inconsistent property quality across datasets are real challenges for automation, schema discovery and data interoperability. In the PG-HIVE framework, the process can be configured to be incrementally processed, so additions in the dataset will indeed have effects to the output. As we haven't tested the incremental option in a real dataset, this would benefit our research, and we can adjust some functionalities to handle your use case — that might be needed in general for real datasets.
As a next step, I’ll run PG-HIVE on the current IYP online dataset, with different clustering parameters, run our evaluation metrics, and provide a first version of the discovered schema.
PG-HIVE captures first node/edge types as a LOOSE SCHEMA and secondly infers data types, constraints as a STRICT SCHEMA.Additionally, when querying the IYP dataset, I found way more node/edge types than the documented. By running PG-HIVE, we can assess how well it captures the details, and you (or another domain expert) can help verify its usefulness or suggest adjustments. This will also give us a sense if the framework needs refinements in bigger and real datasets which will benefit our research. And try to figure out if manual curation is inevitable. Currently we have only tested it in synthetic datasets, with very good results.
In general, I think this will also help your research by providing a concise structure of the dataset, so contributors try to follow this schema when adding the new data. Through this process, you might also find inconsistencies of the dataset, and help the project with normalization or integration of the data.
Reacted by Romain Fontugne and RAJ SHEKHAR
PG Schema
Schema Discovery in property graphs is an emerging concern these days. Real PG datasets lack of a predefined schema (see PG-Schema) that could benefit related tasks, like query optimizations and data integration tasks.
How this project could benefit the community?
As PGs are mostly closed-world datasets (compared to RDF datasets), the community could benefit from this project, if the internet-yellow-page could provide a more precise PG schema, including constraints, data types and cardinalities. The grammar of the PG schema can be found here
As my research focus on PG schema discovery, I could elaborate on this task, using a framework I have implemented (see PG-HIVE), but I need a domain expert to verify the accuracy of the dataset's schema.