Skip to content
This repository was archived by the owner on Aug 12, 2026. It is now read-only.
Merged
Changes from 12 commits
Commits
Show all changes
84 commits
Select commit Hold shift + click to select a range
3ae0d3d
transformation service epic specs
sbilge Oct 9, 2025
dc627fc
transformation service image
sbilge Oct 9, 2025
a360e37
question
sbilge Oct 9, 2025
43fb6b7
removed open question
sbilge Oct 21, 2025
03f79bc
detailed
sbilge Oct 22, 2025
5239e62
minor
sbilge Oct 22, 2025
35e2f95
more specific types
sbilge Oct 22, 2025
b290675
examples added to user journeys
sbilge Oct 24, 2025
72fbd6d
ets revisited
sbilge Nov 17, 2025
ffa654e
deleted outdated image
sbilge Nov 17, 2025
12fe137
collection naming convention
sbilge Nov 18, 2025
64c372f
stricter graph rules
sbilge Nov 18, 2025
a993831
Update 81-question-mark-chrysalis/technical_specification.md
sbilge Nov 18, 2025
a9be559
Update 81-question-mark-chrysalis/technical_specification.md
sbilge Nov 18, 2025
c42aebd
Update 81-question-mark-chrysalis/technical_specification.md
sbilge Nov 18, 2025
9f85310
Update 81-question-mark-chrysalis/technical_specification.md
sbilge Nov 18, 2025
362b677
Update 81-question-mark-chrysalis/technical_specification.md
sbilge Nov 18, 2025
7509573
Update 81-question-mark-chrysalis/technical_specification.md
sbilge Nov 18, 2025
1279660
Update 81-question-mark-chrysalis/technical_specification.md
sbilge Nov 18, 2025
1311b2b
removed mongodb specific terms
sbilge Nov 18, 2025
f3d6d7b
upper case on objects
sbilge Nov 18, 2025
cdfe689
minor Schema model change
sbilge Nov 18, 2025
6afb1ea
minor workflow model change
sbilge Nov 18, 2025
721f3e6
workflow naming schema
sbilge Nov 18, 2025
b5ae4c8
removed type
sbilge Nov 18, 2025
b3006eb
Update 81-question-mark-chrysalis/technical_specification.md
sbilge Nov 18, 2025
7492204
Update 81-question-mark-chrysalis/technical_specification.md
sbilge Nov 18, 2025
2b23a26
Update 81-question-mark-chrysalis/technical_specification.md
sbilge Nov 18, 2025
6f0c0ea
Update 81-question-mark-chrysalis/technical_specification.md
sbilge Nov 18, 2025
c789cd7
Update 81-question-mark-chrysalis/technical_specification.md
sbilge Nov 18, 2025
b101625
Update 81-question-mark-chrysalis/technical_specification.md
sbilge Nov 18, 2025
2269b12
validation steps detailed
sbilge Nov 19, 2025
2ce8a21
Update 81-question-mark-chrysalis/technical_specification.md
sbilge Nov 19, 2025
e1240a1
Update 81-question-mark-chrysalis/technical_specification.md
sbilge Nov 19, 2025
f21dd38
Update 81-question-mark-chrysalis/technical_specification.md
sbilge Nov 19, 2025
26bd754
attribute naming
sbilge Nov 19, 2025
bfa95c9
Update 81-question-mark-chrysalis/technical_specification.md
sbilge Nov 19, 2025
b799edb
obsolete sentence from previous version removed
sbilge Nov 19, 2025
6f4d74f
Update 81-question-mark-chrysalis/technical_specification.md
sbilge Nov 19, 2025
5c89da3
Update 81-question-mark-chrysalis/technical_specification.md
sbilge Nov 19, 2025
2088a69
Update 81-question-mark-chrysalis/technical_specification.md
sbilge Nov 19, 2025
91b802c
Update 81-question-mark-chrysalis/technical_specification.md
sbilge Nov 19, 2025
ac6e5bd
configuration change logic
sbilge Nov 20, 2025
4658982
removed manually populating collections
sbilge Nov 20, 2025
9febd3b
attribute naming
sbilge Nov 20, 2025
5b0e6f6
wording
sbilge Nov 20, 2025
d9ac88f
cross reference
sbilge Nov 20, 2025
a26ddc3
naming, plus configuration logic
sbilge Nov 21, 2025
776be06
edited configs definition
sbilge Nov 21, 2025
0fd0a9f
more explanation on created attribute
sbilge Nov 21, 2025
9c83cde
future extensions are added
sbilge Nov 21, 2025
23cf0e7
remove collections
sbilge Nov 25, 2025
0c3dbb4
Update 81-question-mark-chrysalis/technical_specification.md
sbilge Nov 25, 2025
e3ef19b
Update 81-question-mark-chrysalis/technical_specification.md
sbilge Nov 25, 2025
77905d0
Update 81-question-mark-chrysalis/technical_specification.md
sbilge Nov 25, 2025
7322dad
attribute renamed
sbilge Nov 25, 2025
d952859
Update 81-question-mark-chrysalis/technical_specification.md
sbilge Nov 25, 2025
391f864
Update 81-question-mark-chrysalis/technical_specification.md
sbilge Nov 25, 2025
8d56804
suggested renaming applied to entire text
sbilge Nov 25, 2025
7496fcb
Update 81-question-mark-chrysalis/technical_specification.md
sbilge Nov 25, 2025
ab54e9e
Update 81-question-mark-chrysalis/technical_specification.md
sbilge Nov 25, 2025
f49024e
Update 81-question-mark-chrysalis/technical_specification.md
sbilge Nov 25, 2025
42a4ac4
updated not included section
sbilge Nov 25, 2025
2e6c337
re-wording on validation
sbilge Nov 25, 2025
a2c8255
editing the entire document
sbilge Nov 26, 2025
54d5b3e
minor points added
sbilge Nov 26, 2025
8090173
db in memory
sbilge Nov 26, 2025
385cf2c
Update 81-question-mark-chrysalis/technical_specification.md
sbilge Nov 26, 2025
5c7b67b
Update 81-question-mark-chrysalis/technical_specification.md
sbilge Nov 26, 2025
06d7879
Update 81-question-mark-chrysalis/technical_specification.md
sbilge Nov 26, 2025
c3f6abb
Update 81-question-mark-chrysalis/technical_specification.md
sbilge Nov 26, 2025
8cf5e57
Update 81-question-mark-chrysalis/technical_specification.md
sbilge Nov 26, 2025
855a86d
Update 81-question-mark-chrysalis/technical_specification.md
sbilge Nov 26, 2025
2ae0133
Update 81-question-mark-chrysalis/technical_specification.md
sbilge Nov 27, 2025
f48f066
Model class separation
sbilge Nov 28, 2025
d65312c
Config model renamed
sbilge Nov 28, 2025
6480fa0
Update 81-question-mark-chrysalis/technical_specification.md
sbilge Nov 28, 2025
dc62fe7
merged user journeys
sbilge Nov 28, 2025
fe21c97
BaseModel clarification
sbilge Nov 28, 2025
6728af3
final step on the first journey
sbilge Nov 28, 2025
0c2bb9d
Update 81-question-mark-chrysalis/technical_specification.md
sbilge Nov 28, 2025
127af7e
epic number bump
sbilge Dec 1, 2025
d10cade
index update
sbilge Dec 1, 2025
90ad67e
Merge branch 'main' into question_mark_chrysalis
sbilge Dec 1, 2025
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
230 changes: 230 additions & 0 deletions 81-question-mark-chrysalis/technical_specification.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,230 @@
# Preliminary Experimental Metadata (EM) Transformation Service (Question Mark Chrysalis)
**Epic Type:** Implementation Epic

Epic planning and implementation follow the
[Epic Planning and Marathon SOP](https://ghga.pages.hzdr.de/internal.ghga.de/main/sops/development/epic_planning/).

## Scope

### Outline:

The goal of this epic is to implement a service for the transformation of Experimental Metadata (EM) from one representation/model to another.
This service shall provide functionality around the configurable workflow concept from the `metldata` library to enable these transformations.
To this goal, it needs to keep track of the transformation workflows, original and derived data and workflow routes in its database.


![em-transformation-service overview](./images/em_transformation_service.png)
Comment thread
sbilge marked this conversation as resolved.
Outdated


### Terminology

`Experimental Metadata (EM)`: Data describing the experimental process of generating the Research Data archived in GHGA.

`Experimental Metadata Ingress Models (EMIM)`: Schemas that define the structure and format of experimental metadata as it enters the GHGA system from external sources. Currently, GHGA has a single data ingress model for experimental metadata, which is the ghga-metadata-schema.
Comment thread
sbilge marked this conversation as resolved.
Outdated

`Universal Discovery Model`: A common target representation to which all the EMs are transformed into, used for data queries and data display via th GHGA Data Portal.
Comment thread
sbilge marked this conversation as resolved.
Outdated

`EMPack`: Experimental metadata in datapack format. It follows a relation data schema and compatible with one of the EMIMs.
Comment thread
sbilge marked this conversation as resolved.
Outdated

`AnnotatedEMPack`: A datapack enriched with additional information.
Comment thread
sbilge marked this conversation as resolved.
Outdated

`AnnotatedEMPacks`: A MongoDB collection of AnnotatedEMPack objects.
Comment thread
sbilge marked this conversation as resolved.
Outdated

`Workflow`: Instructions on how to produce a certain data/model representation from another data/model. The workflows are defined by the `metldata` library on which the transformation service will be built.

`Workflows`: A MongoDB collection of Workflow objects.
Comment thread
sbilge marked this conversation as resolved.
Outdated

`WorkflowRoute`: Information on which specific workflow is used to produce data compliant with the output model when presented with data compliant with the input model.

`WorkflowRoutes`: A MongoDB collection of WorkflowRoute objects.
Comment thread
sbilge marked this conversation as resolved.
Outdated

Comment thread
sbilge marked this conversation as resolved.
Outdated

### Included/Required:

- Implement logic to process incoming AnnotatedEMPacks, execute the corresponding workflows and store the resulting AnnotatedEMPacks if required.
- Implement service start-up logic to derive missing schemas for workflow routes with empty output schemas.
Comment thread
sbilge marked this conversation as resolved.
Outdated
- Implement MongoDB collections to store AnnotatedEMPacks, Workflows and WorkflowRoutes.
Comment thread
sbilge marked this conversation as resolved.
Outdated

#### Collections
Comment thread
sbilge marked this conversation as resolved.
Outdated


`Schemas`

- purpose: store schemas

data structure:
```python
class Schema(BaseModel):
id: uuid
name: str
original: bool
version: str | None
content: SchemaPack | None
publish: bool
order: int | None

```

`Workflows`

- Purpose: store metldata-compatible workflow definitions used to transform schemas/data.

data structure:
```python
class TransformationWorkflow(BaseModel):
id: uuid
name: str | None
Comment thread
sbilge marked this conversation as resolved.
Outdated
description: str | None
workflow: metldata.workflow.base.Workflow
```


`WorkflowRoutes`

- Purpose: describe named routes that bind workflows, input schema types and produced output schema types, and whether outputs should be published.
Comment thread
sbilge marked this conversation as resolved.
Outdated

data structure:
```python
class WorkflowRoute(BaseModel):
id: uuid
name: str | None
Comment thread
Cito marked this conversation as resolved.
Outdated
Comment thread
sbilge marked this conversation as resolved.
Outdated
input_schema_id: uuid
output_schema_id: uuid
workflow_id: uuid
```


`AnnotatedEMPacks`

Purpose: store incoming and derived annotated EM datapacks that are to be processed or published.

data structure:
```python
class AnnotatedEMPack(BaseModel):
id: uuid
schema_id: uuid
original_id: uuid | None
datapack: DataPack
annotation: dict
```

#### Implementation Details of Collections
Comment thread
sbilge marked this conversation as resolved.
Outdated

The implementation of the EM transformation service will involve the following key components.
Comment thread
sbilge marked this conversation as resolved.
Outdated

1. `Schemas`

Consists of Schema objects
* `id` is a unique identifier of the schema
* `name` is a human readable name of the schema
* `original` is a boolean flag indicating whether the schema is an original or derived schema
Comment thread
sbilge marked this conversation as resolved.
Outdated
* `version` is the version of the schema (aka type of the schema), initially None if it is not an original schema
Comment thread
sbilge marked this conversation as resolved.
Outdated
* `content` is the schema in schemapack format, initially None if it is not an original schema
Comment thread
sbilge marked this conversation as resolved.
Outdated
* `publish` is a boolean flag indicating whether the derived datapacks conforming to this schema should be published
Comment thread
sbilge marked this conversation as resolved.
Outdated
* `order` is the order in a topological ordering of the schemas in the transformation graph, initially None if it is not an original schema
Comment thread
sbilge marked this conversation as resolved.
Outdated


This collection is populated manually with the defined ghga schemas.

These are original schemas (`original: true`) stored in the `Schemas` and serve as the entry point for data that will be transformed into Universal GHGA Models. EMIM schemas provide a stable interface for data providers and may be more flexible or permissive than internal canonical models to accommodate varied input sources.

2. `Workflows`

Consists of TransformationWorkflow objects
* `id` is a unique identifier for a workflow
* `name` is a human readable name for the workflow
* `description` is a more verbose description of the workflow
Comment thread
sbilge marked this conversation as resolved.
Outdated
* `workflow` is the definition of the workflow in metldata format

This collection is populated manually with the defined ghga transformation workflows.
Comment thread
sbilge marked this conversation as resolved.
Outdated


1. `WorkflowRoutes`
Comment thread
sbilge marked this conversation as resolved.
Outdated

Consists of WorkflowRoute object
* `id` is a unique identifier for the workflow route
* `name` is a human readable name for the workflow route
* `input_schema_id` is the id of the input schema that the workflow route accepts
* `output_schema_id` is the id of the output schema that the workflow route produces
* `workflow_id` is a workflow identifier that is to be applied on a data / schema

Comment thread
sbilge marked this conversation as resolved.
Outdated

This collection is populated manually with the defined ghga workflow routes.
Comment thread
sbilge marked this conversation as resolved.
Outdated
An empty workflow_output_schema indicates the route requires schema derivation at startup.
Comment thread
sbilge marked this conversation as resolved.
Outdated

1. `AnnotatedEMPacks`
Comment thread
sbilge marked this conversation as resolved.
Outdated

Consists of AnnotatedEMPack objects.
Comment thread
sbilge marked this conversation as resolved.
Outdated
* `id` is a unique identifier for the AnnotatedEMPack
* `schema_id` is the identifier of the schema the EM datapack conforms to
* `original_id` is the identifier of the original AnnotatedEMPack from which this AnnotatedEMPack was derived; it is None if this is an original AnnotatedEMPack
* `datapack` is the EM in datapack format that conforms to the schema identified by `schema_id`
* `annotation` is an annotation object with information from other models held by the service

This collection is filled from two sources: incoming events GHGA Study Repository and outputs of executed workflow routes.
Comment thread
sbilge marked this conversation as resolved.
Outdated

Note that we only store published datapacks on the collection. If there are intermediate transformed datapacks, they will only be created and kept in memory as long as they are needed for running a full transformation graph corresponding to a single original datapack.
Comment thread
sbilge marked this conversation as resolved.
Outdated

#### Transformation Configuration

A transformation is configured through the Schemas, Workflows and WorkflowRoutes collections.

Workflow routes defines the transformation paths from input schemas to output schemas using specific workflows. Each workflow route specifies which workflow to apply for a given input schema and what output schema to produce.
Comment thread
sbilge marked this conversation as resolved.
Outdated

Workflow routes also define a a graph structure where schemas are nodes and workflow routes are directed edges. The source nodes of this graph are the “original schemas” aka EM models defined as schemapacks. The graph must not contain any cycles, i.e. it should be a directed acyclic graph (DAG). Additionally, the graph must not contain any "diamonds", i.e. there must be at most one directed path between any two schemas in the graph.
Comment thread
sbilge marked this conversation as resolved.
Outdated

This graph structure is used to validate the transformation configuration and to determine the order in which schemas need to be processed during service startup and when processing incoming AnnotatedEMPacks.

#### Transformation Configuration Validation

Transformation configuration validation includes:

1. Check the graph follows a tree-like structure, i.e. it conforms to the unique path requirement mentioned above.
Comment thread
sbilge marked this conversation as resolved.
Outdated
2. Calculate the topological order of the nodes in the graph using an algorithm similar to Kahn's algorithm, so that the schemas are ordered in a way that if we process the schemas in that order, the previous schemas have already been processed.
3. Update the `order` field of each schema in the `Schemas` according to the calculated topological order.
4. Check the workflows and the original schemas referred by the workflow routes exist in their corresponding collections.
Comment thread
Cito marked this conversation as resolved.
Outdated

#### User Journeys: Schema Derivation (Manual Trigger)

When manually triggered, the service derives the output schemas for all workflow routes with empty output schemas.
Comment thread
sbilge marked this conversation as resolved.
Outdated

1. Validate the transformation configuration of the workflow routes.
Comment thread
sbilge marked this conversation as resolved.
Outdated
2. Traverse the schemas in the transformation graph starting with the original schema, in topological order. For every route execute the following:
1. Retrieve the workflow of the corresponding `workflow_id` from the `Workflows`
2. Retrieve the schema corresponding to the `input_schema_id` from the `WorkflowRoutes`.
3. Call metldata to execute the workflow on the input schema to compute the derived schema.
Comment thread
sbilge marked this conversation as resolved.
Outdated
4. Record the derived schema in the `Schemas` and save its id to the `output_schema_id` of the WorkflowRoute.
Comment thread
sbilge marked this conversation as resolved.
Outdated
3. If there are any errors / conflicts, the operation shall fail and report the error.
Comment thread
sbilge marked this conversation as resolved.
Outdated


#### User Journeys: Service Consumer Receives An Original AnnotatedEMPack
Comment thread
sbilge marked this conversation as resolved.
Outdated

The transformation operation is triggered when the service consumer receives "original annotated Em Pack" or a “re-transform metadata” event.
Comment thread
sbilge marked this conversation as resolved.
Outdated

1. Fetch the datapack ids and corresponding schema ids of all datapacks in the collection where the original_id field matches our original datapack id. Create a mapping of these schema ids to their datapack ids that we call “dirty map”, because all of these datapacks need to be either re-created or deleted.
Comment thread
sbilge marked this conversation as resolved.
Outdated
2. Create a “transformed map” that maps schema ids to datapacks. This will hold the datapacks that have been transformed in this operation already. Initialize it with just the original schemapack id mapping to the original datapack.
3. Traverse the schemas in the transformation graph starting with the original schema, in topological order. For all of these schemas, do the following:
1. Get a route that has the current schema as input. Since we do not allow multiple routes between the same two schemas, there should be at most one such route. Raise an error if there are multiple such routes.
Comment thread
sbilge marked this conversation as resolved.
Outdated
2. Get the datapack from the “transformed map” that corresponds to the input schema of that route. This is the “input datapack” for this step. It should always exist at this point if the topological order was computed correctly, and we can raise an error at this point if this is not the case.
3. Compute the transformed datapack using the input datapack and the workflow route.
4. If the current schema id is in the “dirty map”, remove it there and update the “computed map” with the transformed datapack.
4. Otherwise, add the transformed datapack to the “computed map” with a newly generated id, setting its schema_id to the current schema id, and its original_id field to the original datapack id.
Comment thread
sbilge marked this conversation as resolved.
Outdated
4. Finally, upsert all entries in the “transformed map” to the database that should be published, and delete all remaining entries in the “dirty map” from the database.

This course of action avoids deleting “dirty” datapacks while they are recreated, since that would make the corresponding resources disappear while the transformation is ongoing. It also keeps any intermediate datapacks in memory, avoiding unnecessary database operations and operations.

#### User Journeys: Configuration Change

The “configuration change” operation should be triggered whenever anything in the Schemas, Workflows or WorkflowRoutes collection changes. It will be good to collect such changes and make them all at once, before triggering this operation, because it is very costly.
Comment thread
sbilge marked this conversation as resolved.
Outdated

The configuration change will trigger the following operations:
1. Validate the transformation configuration
2. Recalculate the topological order of the schemas and update the `order` field
3. Re-compute the schemapacks of all transformed schemas
4. Re-run the transformations for all original datapacks

### Not included:

Comment thread
sbilge marked this conversation as resolved.
- The service has no REST API to configure workflows but instead reads workflows from the database, where workflows must be stored in a predefined schema or format as required by the service.
Comment thread
Cito marked this conversation as resolved.
Outdated

- The service has no event consumer to populate its collection of known models but instead reads the models from a database where they have to be stored manually
Comment thread
sbilge marked this conversation as resolved.
Outdated