This repository was archived by the owner on Aug 12, 2026. It is now read-only.
-
Notifications
You must be signed in to change notification settings - Fork 1
Question mark chrysalis #157
Merged
+257
−1
Merged
Changes from 12 commits
Commits
Show all changes
84 commits
Select commit
Hold shift + click to select a range
3ae0d3d
transformation service epic specs
sbilge dc627fc
transformation service image
sbilge a360e37
question
sbilge 43fb6b7
removed open question
sbilge 03f79bc
detailed
sbilge 5239e62
minor
sbilge 35e2f95
more specific types
sbilge b290675
examples added to user journeys
sbilge 72fbd6d
ets revisited
sbilge ffa654e
deleted outdated image
sbilge 12fe137
collection naming convention
sbilge 64c372f
stricter graph rules
sbilge a993831
Update 81-question-mark-chrysalis/technical_specification.md
sbilge a9be559
Update 81-question-mark-chrysalis/technical_specification.md
sbilge c42aebd
Update 81-question-mark-chrysalis/technical_specification.md
sbilge 9f85310
Update 81-question-mark-chrysalis/technical_specification.md
sbilge 362b677
Update 81-question-mark-chrysalis/technical_specification.md
sbilge 7509573
Update 81-question-mark-chrysalis/technical_specification.md
sbilge 1279660
Update 81-question-mark-chrysalis/technical_specification.md
sbilge 1311b2b
removed mongodb specific terms
sbilge f3d6d7b
upper case on objects
sbilge cdfe689
minor Schema model change
sbilge 6afb1ea
minor workflow model change
sbilge 721f3e6
workflow naming schema
sbilge b5ae4c8
removed type
sbilge b3006eb
Update 81-question-mark-chrysalis/technical_specification.md
sbilge 7492204
Update 81-question-mark-chrysalis/technical_specification.md
sbilge 2b23a26
Update 81-question-mark-chrysalis/technical_specification.md
sbilge 6f0c0ea
Update 81-question-mark-chrysalis/technical_specification.md
sbilge c789cd7
Update 81-question-mark-chrysalis/technical_specification.md
sbilge b101625
Update 81-question-mark-chrysalis/technical_specification.md
sbilge 2269b12
validation steps detailed
sbilge 2ce8a21
Update 81-question-mark-chrysalis/technical_specification.md
sbilge e1240a1
Update 81-question-mark-chrysalis/technical_specification.md
sbilge f21dd38
Update 81-question-mark-chrysalis/technical_specification.md
sbilge 26bd754
attribute naming
sbilge bfa95c9
Update 81-question-mark-chrysalis/technical_specification.md
sbilge b799edb
obsolete sentence from previous version removed
sbilge 6f4d74f
Update 81-question-mark-chrysalis/technical_specification.md
sbilge 5c89da3
Update 81-question-mark-chrysalis/technical_specification.md
sbilge 2088a69
Update 81-question-mark-chrysalis/technical_specification.md
sbilge 91b802c
Update 81-question-mark-chrysalis/technical_specification.md
sbilge ac6e5bd
configuration change logic
sbilge 4658982
removed manually populating collections
sbilge 9febd3b
attribute naming
sbilge 5b0e6f6
wording
sbilge d9ac88f
cross reference
sbilge a26ddc3
naming, plus configuration logic
sbilge 776be06
edited configs definition
sbilge 0fd0a9f
more explanation on created attribute
sbilge 9c83cde
future extensions are added
sbilge 23cf0e7
remove collections
sbilge 0c3dbb4
Update 81-question-mark-chrysalis/technical_specification.md
sbilge e3ef19b
Update 81-question-mark-chrysalis/technical_specification.md
sbilge 77905d0
Update 81-question-mark-chrysalis/technical_specification.md
sbilge 7322dad
attribute renamed
sbilge d952859
Update 81-question-mark-chrysalis/technical_specification.md
sbilge 391f864
Update 81-question-mark-chrysalis/technical_specification.md
sbilge 8d56804
suggested renaming applied to entire text
sbilge 7496fcb
Update 81-question-mark-chrysalis/technical_specification.md
sbilge ab54e9e
Update 81-question-mark-chrysalis/technical_specification.md
sbilge f49024e
Update 81-question-mark-chrysalis/technical_specification.md
sbilge 42a4ac4
updated not included section
sbilge 2e6c337
re-wording on validation
sbilge a2c8255
editing the entire document
sbilge 54d5b3e
minor points added
sbilge 8090173
db in memory
sbilge 385cf2c
Update 81-question-mark-chrysalis/technical_specification.md
sbilge 5c7b67b
Update 81-question-mark-chrysalis/technical_specification.md
sbilge 06d7879
Update 81-question-mark-chrysalis/technical_specification.md
sbilge c3f6abb
Update 81-question-mark-chrysalis/technical_specification.md
sbilge 8cf5e57
Update 81-question-mark-chrysalis/technical_specification.md
sbilge 855a86d
Update 81-question-mark-chrysalis/technical_specification.md
sbilge 2ae0133
Update 81-question-mark-chrysalis/technical_specification.md
sbilge f48f066
Model class separation
sbilge d65312c
Config model renamed
sbilge 6480fa0
Update 81-question-mark-chrysalis/technical_specification.md
sbilge dc62fe7
merged user journeys
sbilge fe21c97
BaseModel clarification
sbilge 6728af3
final step on the first journey
sbilge 0c2bb9d
Update 81-question-mark-chrysalis/technical_specification.md
sbilge 127af7e
epic number bump
sbilge d10cade
index update
sbilge 90ad67e
Merge branch 'main' into question_mark_chrysalis
sbilge File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,230 @@ | ||
| # Preliminary Experimental Metadata (EM) Transformation Service (Question Mark Chrysalis) | ||
| **Epic Type:** Implementation Epic | ||
|
|
||
| Epic planning and implementation follow the | ||
| [Epic Planning and Marathon SOP](https://ghga.pages.hzdr.de/internal.ghga.de/main/sops/development/epic_planning/). | ||
|
|
||
| ## Scope | ||
|
|
||
| ### Outline: | ||
|
|
||
| The goal of this epic is to implement a service for the transformation of Experimental Metadata (EM) from one representation/model to another. | ||
| This service shall provide functionality around the configurable workflow concept from the `metldata` library to enable these transformations. | ||
| To this goal, it needs to keep track of the transformation workflows, original and derived data and workflow routes in its database. | ||
|
|
||
|
|
||
|  | ||
|
|
||
|
|
||
| ### Terminology | ||
|
|
||
| `Experimental Metadata (EM)`: Data describing the experimental process of generating the Research Data archived in GHGA. | ||
|
|
||
| `Experimental Metadata Ingress Models (EMIM)`: Schemas that define the structure and format of experimental metadata as it enters the GHGA system from external sources. Currently, GHGA has a single data ingress model for experimental metadata, which is the ghga-metadata-schema. | ||
|
sbilge marked this conversation as resolved.
Outdated
|
||
|
|
||
| `Universal Discovery Model`: A common target representation to which all the EMs are transformed into, used for data queries and data display via th GHGA Data Portal. | ||
|
sbilge marked this conversation as resolved.
Outdated
|
||
|
|
||
| `EMPack`: Experimental metadata in datapack format. It follows a relation data schema and compatible with one of the EMIMs. | ||
|
sbilge marked this conversation as resolved.
Outdated
|
||
|
|
||
| `AnnotatedEMPack`: A datapack enriched with additional information. | ||
|
sbilge marked this conversation as resolved.
Outdated
|
||
|
|
||
| `AnnotatedEMPacks`: A MongoDB collection of AnnotatedEMPack objects. | ||
|
sbilge marked this conversation as resolved.
Outdated
|
||
|
|
||
| `Workflow`: Instructions on how to produce a certain data/model representation from another data/model. The workflows are defined by the `metldata` library on which the transformation service will be built. | ||
|
|
||
| `Workflows`: A MongoDB collection of Workflow objects. | ||
|
sbilge marked this conversation as resolved.
Outdated
|
||
|
|
||
| `WorkflowRoute`: Information on which specific workflow is used to produce data compliant with the output model when presented with data compliant with the input model. | ||
|
|
||
| `WorkflowRoutes`: A MongoDB collection of WorkflowRoute objects. | ||
|
sbilge marked this conversation as resolved.
Outdated
|
||
|
|
||
|
sbilge marked this conversation as resolved.
Outdated
|
||
|
|
||
| ### Included/Required: | ||
|
|
||
| - Implement logic to process incoming AnnotatedEMPacks, execute the corresponding workflows and store the resulting AnnotatedEMPacks if required. | ||
| - Implement service start-up logic to derive missing schemas for workflow routes with empty output schemas. | ||
|
sbilge marked this conversation as resolved.
Outdated
|
||
| - Implement MongoDB collections to store AnnotatedEMPacks, Workflows and WorkflowRoutes. | ||
|
sbilge marked this conversation as resolved.
Outdated
|
||
|
|
||
| #### Collections | ||
|
sbilge marked this conversation as resolved.
Outdated
|
||
|
|
||
|
|
||
| `Schemas` | ||
|
|
||
| - purpose: store schemas | ||
|
|
||
| data structure: | ||
| ```python | ||
| class Schema(BaseModel): | ||
| id: uuid | ||
| name: str | ||
| original: bool | ||
| version: str | None | ||
| content: SchemaPack | None | ||
| publish: bool | ||
| order: int | None | ||
|
|
||
| ``` | ||
|
|
||
| `Workflows` | ||
|
|
||
| - Purpose: store metldata-compatible workflow definitions used to transform schemas/data. | ||
|
|
||
| data structure: | ||
| ```python | ||
| class TransformationWorkflow(BaseModel): | ||
| id: uuid | ||
| name: str | None | ||
|
sbilge marked this conversation as resolved.
Outdated
|
||
| description: str | None | ||
| workflow: metldata.workflow.base.Workflow | ||
| ``` | ||
|
|
||
|
|
||
| `WorkflowRoutes` | ||
|
|
||
| - Purpose: describe named routes that bind workflows, input schema types and produced output schema types, and whether outputs should be published. | ||
|
sbilge marked this conversation as resolved.
Outdated
|
||
|
|
||
| data structure: | ||
| ```python | ||
| class WorkflowRoute(BaseModel): | ||
| id: uuid | ||
| name: str | None | ||
|
Cito marked this conversation as resolved.
Outdated
sbilge marked this conversation as resolved.
Outdated
|
||
| input_schema_id: uuid | ||
| output_schema_id: uuid | ||
| workflow_id: uuid | ||
| ``` | ||
|
|
||
|
|
||
| `AnnotatedEMPacks` | ||
|
|
||
| Purpose: store incoming and derived annotated EM datapacks that are to be processed or published. | ||
|
|
||
| data structure: | ||
| ```python | ||
| class AnnotatedEMPack(BaseModel): | ||
| id: uuid | ||
| schema_id: uuid | ||
| original_id: uuid | None | ||
| datapack: DataPack | ||
| annotation: dict | ||
| ``` | ||
|
|
||
| #### Implementation Details of Collections | ||
|
sbilge marked this conversation as resolved.
Outdated
|
||
|
|
||
| The implementation of the EM transformation service will involve the following key components. | ||
|
sbilge marked this conversation as resolved.
Outdated
|
||
|
|
||
| 1. `Schemas` | ||
|
|
||
| Consists of Schema objects | ||
| * `id` is a unique identifier of the schema | ||
| * `name` is a human readable name of the schema | ||
| * `original` is a boolean flag indicating whether the schema is an original or derived schema | ||
|
sbilge marked this conversation as resolved.
Outdated
|
||
| * `version` is the version of the schema (aka type of the schema), initially None if it is not an original schema | ||
|
sbilge marked this conversation as resolved.
Outdated
|
||
| * `content` is the schema in schemapack format, initially None if it is not an original schema | ||
|
sbilge marked this conversation as resolved.
Outdated
|
||
| * `publish` is a boolean flag indicating whether the derived datapacks conforming to this schema should be published | ||
|
sbilge marked this conversation as resolved.
Outdated
|
||
| * `order` is the order in a topological ordering of the schemas in the transformation graph, initially None if it is not an original schema | ||
|
sbilge marked this conversation as resolved.
Outdated
|
||
|
|
||
|
|
||
| This collection is populated manually with the defined ghga schemas. | ||
|
|
||
| These are original schemas (`original: true`) stored in the `Schemas` and serve as the entry point for data that will be transformed into Universal GHGA Models. EMIM schemas provide a stable interface for data providers and may be more flexible or permissive than internal canonical models to accommodate varied input sources. | ||
|
|
||
| 2. `Workflows` | ||
|
|
||
| Consists of TransformationWorkflow objects | ||
| * `id` is a unique identifier for a workflow | ||
| * `name` is a human readable name for the workflow | ||
| * `description` is a more verbose description of the workflow | ||
|
sbilge marked this conversation as resolved.
Outdated
|
||
| * `workflow` is the definition of the workflow in metldata format | ||
|
|
||
| This collection is populated manually with the defined ghga transformation workflows. | ||
|
sbilge marked this conversation as resolved.
Outdated
|
||
|
|
||
|
|
||
| 1. `WorkflowRoutes` | ||
|
sbilge marked this conversation as resolved.
Outdated
|
||
|
|
||
| Consists of WorkflowRoute object | ||
| * `id` is a unique identifier for the workflow route | ||
| * `name` is a human readable name for the workflow route | ||
| * `input_schema_id` is the id of the input schema that the workflow route accepts | ||
| * `output_schema_id` is the id of the output schema that the workflow route produces | ||
| * `workflow_id` is a workflow identifier that is to be applied on a data / schema | ||
|
|
||
|
sbilge marked this conversation as resolved.
Outdated
|
||
|
|
||
| This collection is populated manually with the defined ghga workflow routes. | ||
|
sbilge marked this conversation as resolved.
Outdated
|
||
| An empty workflow_output_schema indicates the route requires schema derivation at startup. | ||
|
sbilge marked this conversation as resolved.
Outdated
|
||
|
|
||
| 1. `AnnotatedEMPacks` | ||
|
sbilge marked this conversation as resolved.
Outdated
|
||
|
|
||
| Consists of AnnotatedEMPack objects. | ||
|
sbilge marked this conversation as resolved.
Outdated
|
||
| * `id` is a unique identifier for the AnnotatedEMPack | ||
| * `schema_id` is the identifier of the schema the EM datapack conforms to | ||
| * `original_id` is the identifier of the original AnnotatedEMPack from which this AnnotatedEMPack was derived; it is None if this is an original AnnotatedEMPack | ||
| * `datapack` is the EM in datapack format that conforms to the schema identified by `schema_id` | ||
| * `annotation` is an annotation object with information from other models held by the service | ||
|
|
||
| This collection is filled from two sources: incoming events GHGA Study Repository and outputs of executed workflow routes. | ||
|
sbilge marked this conversation as resolved.
Outdated
|
||
|
|
||
| Note that we only store published datapacks on the collection. If there are intermediate transformed datapacks, they will only be created and kept in memory as long as they are needed for running a full transformation graph corresponding to a single original datapack. | ||
|
sbilge marked this conversation as resolved.
Outdated
|
||
|
|
||
| #### Transformation Configuration | ||
|
|
||
| A transformation is configured through the Schemas, Workflows and WorkflowRoutes collections. | ||
|
|
||
| Workflow routes defines the transformation paths from input schemas to output schemas using specific workflows. Each workflow route specifies which workflow to apply for a given input schema and what output schema to produce. | ||
|
sbilge marked this conversation as resolved.
Outdated
|
||
|
|
||
| Workflow routes also define a a graph structure where schemas are nodes and workflow routes are directed edges. The source nodes of this graph are the “original schemas” aka EM models defined as schemapacks. The graph must not contain any cycles, i.e. it should be a directed acyclic graph (DAG). Additionally, the graph must not contain any "diamonds", i.e. there must be at most one directed path between any two schemas in the graph. | ||
|
sbilge marked this conversation as resolved.
Outdated
|
||
|
|
||
| This graph structure is used to validate the transformation configuration and to determine the order in which schemas need to be processed during service startup and when processing incoming AnnotatedEMPacks. | ||
|
|
||
| #### Transformation Configuration Validation | ||
|
|
||
| Transformation configuration validation includes: | ||
|
|
||
| 1. Check the graph follows a tree-like structure, i.e. it conforms to the unique path requirement mentioned above. | ||
|
sbilge marked this conversation as resolved.
Outdated
|
||
| 2. Calculate the topological order of the nodes in the graph using an algorithm similar to Kahn's algorithm, so that the schemas are ordered in a way that if we process the schemas in that order, the previous schemas have already been processed. | ||
| 3. Update the `order` field of each schema in the `Schemas` according to the calculated topological order. | ||
| 4. Check the workflows and the original schemas referred by the workflow routes exist in their corresponding collections. | ||
|
Cito marked this conversation as resolved.
Outdated
|
||
|
|
||
| #### User Journeys: Schema Derivation (Manual Trigger) | ||
|
|
||
| When manually triggered, the service derives the output schemas for all workflow routes with empty output schemas. | ||
|
sbilge marked this conversation as resolved.
Outdated
|
||
|
|
||
| 1. Validate the transformation configuration of the workflow routes. | ||
|
sbilge marked this conversation as resolved.
Outdated
|
||
| 2. Traverse the schemas in the transformation graph starting with the original schema, in topological order. For every route execute the following: | ||
| 1. Retrieve the workflow of the corresponding `workflow_id` from the `Workflows` | ||
| 2. Retrieve the schema corresponding to the `input_schema_id` from the `WorkflowRoutes`. | ||
| 3. Call metldata to execute the workflow on the input schema to compute the derived schema. | ||
|
sbilge marked this conversation as resolved.
Outdated
|
||
| 4. Record the derived schema in the `Schemas` and save its id to the `output_schema_id` of the WorkflowRoute. | ||
|
sbilge marked this conversation as resolved.
Outdated
|
||
| 3. If there are any errors / conflicts, the operation shall fail and report the error. | ||
|
sbilge marked this conversation as resolved.
Outdated
|
||
|
|
||
|
|
||
| #### User Journeys: Service Consumer Receives An Original AnnotatedEMPack | ||
|
sbilge marked this conversation as resolved.
Outdated
|
||
|
|
||
| The transformation operation is triggered when the service consumer receives "original annotated Em Pack" or a “re-transform metadata” event. | ||
|
sbilge marked this conversation as resolved.
Outdated
|
||
|
|
||
| 1. Fetch the datapack ids and corresponding schema ids of all datapacks in the collection where the original_id field matches our original datapack id. Create a mapping of these schema ids to their datapack ids that we call “dirty map”, because all of these datapacks need to be either re-created or deleted. | ||
|
sbilge marked this conversation as resolved.
Outdated
|
||
| 2. Create a “transformed map” that maps schema ids to datapacks. This will hold the datapacks that have been transformed in this operation already. Initialize it with just the original schemapack id mapping to the original datapack. | ||
| 3. Traverse the schemas in the transformation graph starting with the original schema, in topological order. For all of these schemas, do the following: | ||
| 1. Get a route that has the current schema as input. Since we do not allow multiple routes between the same two schemas, there should be at most one such route. Raise an error if there are multiple such routes. | ||
|
sbilge marked this conversation as resolved.
Outdated
|
||
| 2. Get the datapack from the “transformed map” that corresponds to the input schema of that route. This is the “input datapack” for this step. It should always exist at this point if the topological order was computed correctly, and we can raise an error at this point if this is not the case. | ||
| 3. Compute the transformed datapack using the input datapack and the workflow route. | ||
| 4. If the current schema id is in the “dirty map”, remove it there and update the “computed map” with the transformed datapack. | ||
| 4. Otherwise, add the transformed datapack to the “computed map” with a newly generated id, setting its schema_id to the current schema id, and its original_id field to the original datapack id. | ||
|
sbilge marked this conversation as resolved.
Outdated
|
||
| 4. Finally, upsert all entries in the “transformed map” to the database that should be published, and delete all remaining entries in the “dirty map” from the database. | ||
|
|
||
| This course of action avoids deleting “dirty” datapacks while they are recreated, since that would make the corresponding resources disappear while the transformation is ongoing. It also keeps any intermediate datapacks in memory, avoiding unnecessary database operations and operations. | ||
|
|
||
| #### User Journeys: Configuration Change | ||
|
|
||
| The “configuration change” operation should be triggered whenever anything in the Schemas, Workflows or WorkflowRoutes collection changes. It will be good to collect such changes and make them all at once, before triggering this operation, because it is very costly. | ||
|
sbilge marked this conversation as resolved.
Outdated
|
||
|
|
||
| The configuration change will trigger the following operations: | ||
| 1. Validate the transformation configuration | ||
| 2. Recalculate the topological order of the schemas and update the `order` field | ||
| 3. Re-compute the schemapacks of all transformed schemas | ||
| 4. Re-run the transformations for all original datapacks | ||
|
|
||
| ### Not included: | ||
|
|
||
|
sbilge marked this conversation as resolved.
|
||
| - The service has no REST API to configure workflows but instead reads workflows from the database, where workflows must be stored in a predefined schema or format as required by the service. | ||
|
Cito marked this conversation as resolved.
Outdated
|
||
|
|
||
| - The service has no event consumer to populate its collection of known models but instead reads the models from a database where they have to be stored manually | ||
|
sbilge marked this conversation as resolved.
Outdated
|
||
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
Uh oh!
There was an error while loading. Please reload this page.