feat(backend): postgres integration - #12379
Conversation
|
Hi @kaikaila. Thanks for your PR. I'm waiting for a kubeflow member to verify that this patch is reasonable to test. If it is, they should reply with Once the patch is verified, the new status will be reflected by the I understand the commands that are listed here. DetailsInstructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes/test-infra repository. |
|
🚫 This command cannot be processed. Only organization members or owners can use the commands. |
cd1d08b to
85498ed
Compare
|
Currently, both MySQL and PGX setups use the DB superuser for all KFP operations, which is why client_manager.go contains a “create database if not exist” step here. From a security standpoint, would it be preferable to:
If the team agrees, I can propose a follow-up PR to refactor accordingly. |
|
I'm fine with this, I don't think it's great that KFP tries to create a database (or a bucket frankly) fyi @mprahl / @droctothorpe |
|
Thanks, @HumairAK — totally agree on the security point. |
09fd370 to
1e0caa8
Compare
|
yes that is fine |
4d33821 to
e6c943c
Compare
Question about the PostgreSQL test workflow organizationCurrent situationThe V2 integration tests for PostgreSQL logically belong in a "PostgreSQL counterpart" to legacy-v2-api-integration-tests.yml Question: What's the recommended workflow organization for PostgreSQL tests?Should I:
Would love guidance on the long-term vision for test workflow organization, especially from @nsingla |
| fi | ||
|
|
||
| # Manifests will be deployed according to the flag provided | ||
| # Manifest selection: each branch picks ONE pre-built kustomize overlay directory. |
There was a problem hiding this comment.
Hi @HumairAK, regarding the if-else chain in deploy-kfp.sh (lines 144-221) — I've been thinking about how to improve it but haven't landed on a good approach.
The chain is ugly, but it maps directly to the CI test matrix combinations. Since this script is only used by CI, refactoring it into a compositional "building blocks" approach has limited ROI. More importantly, it's not straightforward to do — options like tls-enabled and postgresql each use a different kustomize base (platform-agnostic-standalone-tls vs platform-agnostic-postgresql), so they can't simply be stacked as independent overlays.
For now I've added a comment block before the if-else chain explaining why it's mutually exclusive priority matching rather than free combination of flags, so future readers aren't confused by the structure.
Do you have any suggestions for improving this, or are you okay leaving it as-is with the added comments?
| // quoteDSNValue quotes a value for use in a libpq keyword/value connection string. | ||
| // Per libpq rules: wrap in single quotes, escape \ and ' with a backslash. | ||
| func quoteDSNValue(v string) string { | ||
| v = strings.ReplaceAll(v, `\`, `\\`) | ||
| v = strings.ReplaceAll(v, `'`, `\'`) | ||
| return "'" + v + "'" | ||
| } |
There was a problem hiding this comment.
The quoting in dialect.go is for SQL queries — QuoteIdentifier escapes SQL identifiers (table/column names), and EscapeSQLString escapes SQL string literals. These are meant to be sent to the database engine for execution.
The quoting in the DSN is for the libpq connection string — it follows libpq's own parsing rules, which have nothing to do with SQL syntax. libpq's rule is: wrap values in single quotes, and escape ' and \ within values using a \ prefix.
These are two fundamentally different escaping schemes, so the dialect package methods should not be reused here. The DSN value escaping belongs in its own small helper, co-located in sql.go.
| // Wait a bit more to ensure the first run's launcher has finished writing cache to database | ||
| // The run state becomes SUCCEEDED when the user container finishes, but launcher still needs | ||
| // time to publish results and create cache entry in the database | ||
| time.Sleep(15 * time.Second) | ||
|
|
||
| state := s.getContainerExecutionState(t, allRuns[1].RunID) |
There was a problem hiding this comment.
This suggestion has limited value in this scenario — the real timing risk isn't in this 15-second sleep, but in whether Run#1's cache entry has been written to the database by the time Run#2 starts. That timing is controlled by the recurring run's 60-second interval, not by the test code, so polling here wouldn't address the actual race condition.
| tests := []struct { | ||
| name string | ||
| tasks []*model.Task | ||
| runID string | ||
| want []*model.Task | ||
| wantErr bool | ||
| errMsg string |
There was a problem hiding this comment.
If the field is deleted, there will be many compiler errors. Instead, I replaced defaultFakeRunIdTwo with tt.runID in line 514.
| // Also check runtime value type as fallback for old tokens that lack this field. | ||
| _, valueIsString := o.SortByFieldValue.(string) | ||
| if o.SortByFieldIsString || valueIsString { | ||
| sqlBuilder = sqlBuilder.OrderBy(fmt.Sprintf("LOWER(%v) %v", sortByFieldNameWithPrefix, order)) |
There was a problem hiding this comment.
It is intentionally put in the shared layer to keep case-insensitive behavior across all resources, not just pipelines. Following the decision here.
| fi | ||
|
|
||
| # Manifests will be deployed according to the flag provided | ||
| # Manifest selection: each branch picks ONE pre-built kustomize overlay directory. |
There was a problem hiding this comment.
Good catch, thanks! I looked into pinning to a release tag (e.g. the latest v1.10.2 is released in July 2025 ), but it turns out the profile-controller path (applications/dashboard/upstream/profile-controller/...) was only added to kubeflow/manifests master in April 2026 and has never been included in any release tag. So pinning to a tag would break that path with a 404.
Instead, I've pinned all four kubectl apply -k calls to a specific commit SHA (June 2026), which gives us reproducibility without requiring a tagged release.
| assert.Equal(t, 5, totalSize) | ||
| assert.Equal(t, "arguments-parameters.yaml", listFirstPagePipelines[1].Name) | ||
| assert.Equal(t, "arguments_parameters.zip", listFirstPagePipelines[0].Name) | ||
| // MySQL: _ sorts before - (ascending), so zip comes first at index 0 |
There was a problem hiding this comment.
Let me check with @nsingla — you introduced these fixture names in #12440. Was there a specific reason we needed arguments-parameters.yaml / arguments_parameters.zip as the sort sentinels, or is it safe to rename them to alphanumeric-only names to avoid collation sensitivity? If there's no hard dependency, I'll swap them out for simpler names.
| if s, ok := v.(string); ok { | ||
| col := QualifyIdentifier(quote, k) | ||
| andExprs = append(andExprs, squirrel.Expr( | ||
| fmt.Sprintf("LOWER(%s) = LOWER(?)", col), s, |
There was a problem hiding this comment.
Tracked separately in #13512, which makes matchesFilter() case-insensitive to match the SQL-backed LOWER() behavior. Let's continue the discussion there.
| num_parallel_nodes: ${{ env.NUMBER_OF_PARALLEL_NODES }} | ||
| default_namespace: ${{ env.NAMESPACE }} | ||
| python_version: ${{ env.PYTHON_VERSION }} | ||
| report_name: "K8Native_k8sVersion=${{ matrix.k8s_version }}_cacheEnabled=${{ matrix.cache_enabled }}_argoVersion=${{ matrix.argo_version }}_uploadPipelinesWithKubernetesClient=${{ matrix.uploadPipelinesWithKubernetesClient }}" |
There was a problem hiding this comment.
Hi @mprahl Good Catch. Though, these are pre-existing issues on master, which are out of scope of this PR. I've raised them separately in #13594.
Separately, the k8s-native job currently has no db_type axis, but it still relies on a database for storing runs, experiments, jobs, and other non-pipeline resources. Should we add an explicit db_type axis to this job (starting with MySQL only, and expanding to PostgreSQL in a follow-up)?
| psql -h 127.0.0.1 -p 5432 -U user -d mlpipeline | ||
| ``` | ||
|
|
||
| When prompted for a password, enter: `password` |
There was a problem hiding this comment.
For context, @HumairAK and I discussed the least-privilege refactor earlier (proposal — the plan is to move CREATE DATABASE out of client_manager.go into the deployment/init phase and introduce a dedicated restricted user in a follow-up PR.
I think the docs change you're requesting here should stay in sync with that code change. If we only update the docs to instruct users to create a least-privileged user now, the API server would fail at CREATE DATABASE which requires superuser privileges.
So I suggest deferring both docs & code land together in the follow-up PR, tracking in #13797
| var sortByField interface{} | ||
| if sortByField = listable.GetFieldValue(o.SortByFieldName); sortByField == nil { | ||
| return nil, util.NewInvalidInputError("cannot sort by field %q on type %q", o.SortByFieldName, elemName) | ||
| if sortByField = listable.GetFieldValue(fieldNameForValue); sortByField == nil { |
There was a problem hiding this comment.
GPT 5.6 Sol Review
When sorting by a metric, the SQL deliberately produces NULL for runs without that metric, but this path rejects the lookahead row while generating the next-page token. If the first omitted run has no matching metric, an otherwise valid ListRuns request fails with cannot sort by field instead of returning a token. PostgreSQL and MySQL also place NULLs differently by default, so this varies by backend and direction. Could we define deterministic NULL ordering, represent NULL in the token, and add multi-page tests containing runs both with and without the selected metric?
| operation = func() error { | ||
| _, err = db.Exec(fmt.Sprintf("CREATE DATABASE %s", dbName)) | ||
| if ignoreAlreadyExistError(dialect, err) != nil { | ||
| _, err = db.Exec(fmt.Sprintf("CREATE DATABASE %s", drvDialect.QuoteIdentifier(dbName))) |
There was a problem hiding this comment.
GPT 5.6 Sol Review
This unconditionally executes CREATE DATABASE even when the target database has already been provisioned. A normal runtime role that owns the existing database but lacks CREATEDB receives permission denied to create database; the cache server has the same behavior. This forces the bundled deployment to give KFP a PostgreSQL superuser and prevents common managed/least-privilege configurations. Could database creation move to an initialization/admin phase, or could we support a pre-provisioned-database mode so runtime components can use restricted roles?
| extraParams map[string]string, | ||
| ) (*pgx.ConnConfig, string, error) { | ||
| q := url.Values{} | ||
| q.Set("sslmode", "disable") |
There was a problem hiding this comment.
GPT 5.6 Sol Review
The new PostgreSQL client still defaults to sslmode=disable, so production-facing deployments send database traffic without transport encryption unless operators discover and override the extra parameter. The cache helper has the same default, and the new overlays do not set a safer value. Could the development manifest opt into disable explicitly while the client uses a secure default (or requires an explicit SSL mode), with documented verify-full support for production?
There was a problem hiding this comment.
Done for postgres: require sslmode to be set explicitly instead: the function now returns an error (fail-fast / CrashLoop) when sslmode is absent, forcing a deliberate choice ("disable" for local dev, "verify-full" for production).
Open #13796 to track the parity feature for MySQL.
| operation = func() error { | ||
| _, err = db.Exec(fmt.Sprintf("CREATE DATABASE IF NOT EXISTS %s", dbName)) | ||
| var pgxExtraParams = map[string]string{} | ||
| json.Unmarshal([]byte(params.dbExtraParams), &pgxExtraParams) |
There was a problem hiding this comment.
GPT 5.6 Sol Review
The JSON parse error is discarded here. A typo in --db_extra_params, including intended TLS options such as sslmode=verify-full, silently produces an empty map and falls back to sslmode=disable; the operator receives no indication that the requested security configuration was ignored. Could startup fail on malformed JSON and cover malformed PostgreSQL and MySQL parameter maps in tests?
| for _, v := range vs { | ||
| if s, ok := v.(string); ok { | ||
| col := QualifyIdentifier(quote, k) | ||
| andExprs = append(andExprs, squirrel.Expr( |
There was a problem hiding this comment.
GPT 5.6 Sol Review
Applying LOWER() to every string predicate changes exact-match semantics for identifiers and enums such as UUIDs, namespaces, and storage states. On PostgreSQL it can also prevent ordinary indexes and primary-key indexes from serving those filters. User-facing fields such as display names may reasonably be case-insensitive, but identifiers should retain direct comparison. Could we whitelist fields intended to be case-insensitive and preserve exact comparisons for identifiers/enums?
juliusvonkohout
left a comment
There was a problem hiding this comment.
One problem i saw recently was a bit of manifest duplication. E.g we now have 2 pipeline-install configmaps in the repository instead of patching the existing one. So please make sure that you do not duplicate manifests and reduce existing duplication for postgres.
Hi @juliusvonkohout Good catch! I removed the file in the latest push. |
| argo_version: [ "v3.7.14", "v4.0.5" ] | ||
| pipeline_store: [ "database" ] | ||
| pod_to_pod_tls_enabled: [ "false" ] | ||
| db_type: ["mysql", "pgx"] |
There was a problem hiding this comment.
We should add this option to the e2e-tests.yml as well
There was a problem hiding this comment.
Upgrade tests should also have this option as well
There was a problem hiding this comment.
Added db_type: "pgx" to all three E2E jobs in e2e-test.yml using include entries (one representative combination per job).
The current approach uses include enumeration rather than a top-level matrix dimension because pgx kustomize overlays don't exist for every dimension combination yet:
Standalone: pgx is incompatible with cache_disabled, proxy, and pod_to_pod_tls (no corresponding overlays in standalone/postgresql/)
Multi-user: pgx is incompatible with both cache_disabled and artifact_proxy (no overlays in multiuser/postgresql/), so the multi-user pgx entry must explicitly disable artifact_proxy
Each pgx include has a TODO comment noting which overlays would need to be created to expand coverage. If broader coverage is preferred for this PR (e.g., promoting db_type to a top-level dimension with excludes for the standalone job, similar to api-server-tests.yml), happy to expand — that would add ~8 runners for Job 1. Alternatively, we can expand incrementally in follow-up PRs as more overlays are added and keep a tracking issue.
There was a problem hiding this comment.
The pgx option to upgrade test doesn't make sense yet: since no release has ever successfully run with PostgreSQL, there is no pre-existing PostgreSQL data to migrate.
Plan: Once this PR merges and a release with full PostgreSQL support is published (expected 2.18), there will be a real upgrade path to validate. I'll file a tracking issue to make sure this doesn't fall through the cracks.
Fresh pgx deployment and functionality are already covered by e2e-test.yml, api-server-tests.yml, and integration-tests-v1.yml (which also skips its upgrade test for pgx via -skip TestUpgrade).
| cache_enabled: "true" | ||
| pod_to_pod_tls_enabled: "true" | ||
| db_type: "mysql" | ||
| exclude: |
There was a problem hiding this comment.
why are these excluded? is this an incompatibility that's documented?
There was a problem hiding this comment.
These combinations are excluded because the corresponding CI kustomize overlays don't exist yet. Specifically:
- pgx + proxy=true — no standalone/postgresql/proxy overlay
- pgx + cache_enabled=false — no standalone/postgresql/cache-disabled overlay
The deploy script (deploy-kfp.sh) also has explicit guard-rails that fail fast if someone tries to run these unsupported combinations, so this isn't a silent gap.
For this PR I intentionally kept the PostgreSQL CI coverage to the core path (cache_enabled=true, proxy=false) to validate the integration end-to-end without expanding the overlay surface all at once. The multi-user section has a TODO comment noting the same (missing pgx + cache_disabled entry).
Happy to create a follow-up issue to track adding the remaining overlay variants (proxy, cache-disabled, TLS) — would that work?
There was a problem hiding this comment.
My concern with that is that until its tested, we cannot be certain that its a supported path, unless you've tested it manually to confirm that
There was a problem hiding this comment.
You're right that an untested path shouldn't be presented as supported. I've documented it explicitly on the operator-facing side: kubeflow/website#4438. The website PR is written against the state after this PR merges to master, so I'll hold it until then.
Since no multi-user pod-to-pod TLS overlay exists for any database backend, I also raised an issue here.
| strategy: | ||
| matrix: | ||
| k8s_version: [ "v1.36.1"] | ||
| k8s_version: [ "v1.36.1" ] |
There was a problem hiding this comment.
we should add db_type here as well
Review GuideThis PR touches ~150 files. To make review manageable, changes are grouped by module below. Recommended review order: start with the dialect abstraction (the foundation), then stores, then tests, then deployment. 1. Dialect Abstraction Layer (start here)Core abstraction that all other changes depend on.
Deleted (replaced by the above):
2. Store Layer UpdatesEach store updated to accept
3. Cross-Database Behavioral DifferencesMySQL and PostgreSQL differ in three areas that required explicit handling (see PR description for details):
4. Client Manager & Cache ServiceInitialization and cache service plumbing for dual-DB support.
Deleted: 5. Test Infrastructure & Integration TestsTest stability fixes triggered by postgres's different runtime behavior (stricter collation, different timing).
6. Deployment Manifests & CICan be reviewed independently from Go code.
Deleted (replaced by new structure):
7. Documentation
|
| if pluginsInput.Valid { | ||
| lt := model.LargeText(pluginsInput.String) | ||
| run.PluginsInputString = < | ||
| } | ||
| if pluginsOutput.Valid { | ||
| lt := model.LargeText(pluginsOutput.String) | ||
| run.PluginsOutputString = < | ||
| } |
There was a problem hiding this comment.
This duplicates the above lines.
| if pluginsInput.Valid { | |
| lt := model.LargeText(pluginsInput.String) | |
| run.PluginsInputString = < | |
| } | |
| if pluginsOutput.Valid { | |
| lt := model.LargeText(pluginsOutput.String) | |
| run.PluginsOutputString = < | |
| } |
There was a problem hiding this comment.
The PostgreSQL CI overlays (both standalone and multiuser) are missing the ../../base/ci-stability-tuning.yaml patch that all other CI overlays include. This patch increases probe timeouts and failure thresholds to prevent flaky CI failures under load. Without it, PostgreSQL CI jobs may fail intermittently when the API server takes longer to start.
Verified: standalone/default/kustomization.yaml includes - path: ../../base/ci-stability-tuning.yaml but standalone/postgresql/kustomization.yaml does not. Same for multiuser/postgresql.
Suggested fix: Add to both standalone/postgresql/kustomization.yaml and multiuser/postgresql/kustomization.yaml under patches::
- path: ../../base/ci-stability-tuning.yaml|
|
||
| // DBDialect holds read-only runtime configuration for a SQL backend. | ||
| // All fields are private; callers must use the exported getter methods. | ||
| type DBDialect struct { |
There was a problem hiding this comment.
DBDialect should be an interface, not a concrete struct. The codebase already had a SQLDialect interface in backend/src/apiserver/storage/db.go that used method dispatch for dialect-specific behavior. This PR replaced it with a concrete struct that uses switch d.name blocks and stores QuoteIdentifier as a lambda — a regression from the existing design.
Currently DBDialect uses switch d.name blocks inside ConcatAgg and ConcatExprs, each with a default: panic(...) branch. The QuoteFunction type alias exists only because quoting is stored as a function pointer rather than being a method.
With an interface + per-dialect implementations (mysqlDialect, pgxDialect, sqliteDialect):
type DBDialect interface {
Name() string
QuoteIdentifier(string) string
LengthFunc() string
QueryBuilder() sq.StatementBuilderType
ExistDatabaseErrHint() string
StringCollation() string
ConcatAgg(distinct bool, expr, sep string) string
ConcatExprs(exprs []string, sep string) string
}This eliminates:
- 2
switch d.nameblocks (replaced by method dispatch) - 2
default: panic(...)branches in methods (the compiler enforces every dialect implements every method) - The
QuoteFunctiontype alias (quoting becomes a regular method)
isDuplicateError and insertUpsert in sql_dialect_util.go could also become interface methods, which would isolate driver-specific imports (go-sql-driver/mysql, pgconn, go-sqlite3) into each concrete type rather than concentrating them in one file.
| sqlBuilder = opts.AddOrderByToSelect(sqlBuilder, q, s.dbDialect.StringCollation()) | ||
| } | ||
| // Apply correct placeholder format for the final SQL generation | ||
| if s.dbDialect.Name() == "pgx" { |
There was a problem hiding this comment.
Four locations manually check s.dbDialect.Name() == "pgx" to re-apply PlaceholderFormat(sq.Dollar) after subquery construction: run_store.go:228, run_store.go:248, job_store.go:185, job_store.go:201. This is a dialect abstraction leak — stores should not inspect the dialect name.
Suggested fix: Add a helper to DBDialect (e.g., FinalizeSelect(sq.SelectBuilder) sq.SelectBuilder) that applies the correct final placeholder format. Replace all four string-comparison checks with the helper call.
| // Otherwise, return nil. | ||
| func ignoreAlreadyExistError(dialect SQLDialect, err error) error { | ||
| if err != nil && strings.Contains(err.Error(), dialect.ExistDatabaseErrHint) { | ||
| func ignoreAlreadyExistError(dialect sqldrv.DBDialect, err error) error { |
There was a problem hiding this comment.
ignoreAlreadyExistError uses strings.Contains(err.Error(), "already exists") for PostgreSQL. The PR already uses structured pgerrcode checking in isDuplicateError (sql_dialect_util.go). This function should use pgerrcode.DuplicateDatabase for consistency and robustness.
Suggested fix: Use errors.As with *pgconn.PgError and check pe.Code == pgerrcode.DuplicateDatabase for the pgx case (both packages are already in go.mod).
|
|
||
| ```go | ||
| func (s *ExperimentStore) ArchiveExperiment(id string) error { | ||
| quotedTable := s.dialect.QuoteIdentifier("Experiments") |
There was a problem hiding this comment.
The Squirrel example in the "Database query guidelines" section has multiple errors that would produce broken code on PostgreSQL if copied:
s.dialect.QuoteIdentifier("Experiments")— table name is"experiments"(lowercase), and the field iss.dbDialect, nots.dialect.Set(quotedState, ...)— should be.SetMap(sq.Eq{...})
The GORM example at line 703 uses Where("uuid = ?", uuid) with a raw lowercase column name. GORM maps struct fields through tags so this would actually use the tagged column "UUID", but the example is misleading — a reader might think raw lowercase column names work in Squirrel queries too.
|
|
||
| ### Database Schema Naming Convention | ||
|
|
||
| **Important**: KFP uses **CamelCase** table and column names (e.g., `Experiments`, `ExperimentUUID`) as a legacy design choice from the original MySQL implementation. |
There was a problem hiding this comment.
"KFP uses CamelCase table and column names (e.g., Experiments, ExperimentUUID)" — table names are actually lowercase (experiments, pipelines, run_details, pipeline_versions). Only column names are CamelCase. This distinction matters because a developer quoting "Experiments" on PostgreSQL will get a "relation does not exist" error.
There was a problem hiding this comment.
The production PostgreSQL manifests (platform-agnostic-postgresql and platform-agnostic-multi-user-postgresql) don't set DBCONFIG_POSTGRESQLCONFIG_EXTRAPARAMS with sslmode. The code at common/sql/config.go:60 requires sslmode explicitly and crashes without it. Anyone deploying with kubectl apply -k manifests/kustomize/env/platform-agnostic-postgresql — as documented in the PR's Migration Guide — will get a CrashLoopBackOff.
CI passes because the CI overlays (.github/resources/manifests/standalone/postgresql/apiserver-env.yaml) add sslmode=disable via their own patches. The cache server has the same issue — cache-server-patch.yaml doesn't pass --db_extra_params='{"sslmode":"disable"}', so it will also crash.
Suggested fix: Set sslmode=disable in the production manifests with a comment to override for production TLS — the same approach MySQL uses (ships with working defaults, documents what to change for production). The current "secure-by-default" design that crashes instead of starting is not secure, it's broken — no user can deploy PostgreSQL without first discovering they need an undocumented env var override.
There was a problem hiding this comment.
This sounds good to me. We can create a follow up GitHub issue which adds manifests for KFP with Postgres with TLS enabled and using cert-manager to provision the certs.
| test_label: "E2ECritical" | ||
| # PostgreSQL lane — cache_enabled=true only (no pgx overlay for | ||
| # cache-disabled, proxy, or TLS). | ||
| # TODO: expand when additional standalone/postgresql/* overlays exist. |
There was a problem hiding this comment.
Why don't they exist? Shouldn't this be part of this PR?
There was a problem hiding this comment.
The initial PostgreSQL scope intentionally covers only the core deployment configuration. The missing combinations (cache-disabled, proxy, pod-to-pod TLS, and multi-user artifact proxy) require dedicated PostgreSQL overlays and CI coverage.
They are tracked in follow-up issue #13822, which is also listed in this PR’s description.
| root_password: password No newline at end of file | ||
| stringData: | ||
| # This user is used by KFP components (like apiserver) to connect to the database. | ||
| # TODO(kaikaila): make it a normal user instead of a superuser. |
There was a problem hiding this comment.
No, it’s intentional and tracked as a follow-up in #13797.
commented
Aug 10, 2026
|
@hsinhoyeh for collaboration and #13956 |
commented
Aug 21, 2026
@kaikaila Just checking in on this PR—what's the current status? |
commented
Aug 22, 2026
|
@hsinatfootprintai in progress. Thanks for your patience |
commented
Aug 27, 2026
|
Hi @hbelmiro — PR #12379 is ready for another review. I addressed the comments from your previous review round. I split the current branch into two commits to keep the review focused:
Key updates:
Thanks for taking another look. |
Signed-off-by: kaikaila <lyk2772@126.com> Signed-off-by: Yunkai Li <yunli@redhat.com>
commented
Sep 8, 2026
|
[APPROVALNOTIFIER] This PR is NOT APPROVED This pull-request has been approved by: hbelmiro The full list of commands accepted by this bot can be found here. DetailsNeeds approval from an approver in each of these files:Approvers can indicate their approval by writing |
Summary
This PR adds full PostgreSQL (pgx driver) support to Kubeflow Pipelines backend, enabling users to choose between MySQL and PostgreSQL as the metadata database. The implementation introduces a clean dialect abstraction layer and includes a major query optimization that benefits both database backends.
Fixes #7512
Fixes #9813
Key achievements:
the root causes behind [frontend] Latency when listing runs #10778, [frontend] Cannot list runs and/or artifacts (upstream request timeout) #10230, [backend] performance issues with list runs API #9780, [frontend] Listing all runs take a lot of time #9701
What Changed
1. Storage Layer Refactoring - Dialect Abstraction
DBDialectinterface encapsulating database-specific identifier quoting, placeholders, and aggregation.backend/src/apiserver/storage/list_filters.go).2. ListRuns Query Performance Optimization
3. Deployment & CI Configurations
platform-agnostic-postgresql.make DATABASE=postgres dev-kind-cluster).db_type: ["mysql", "pgx"].4. Consistency & Inconsitency across backend databases
NULLordering for sorts across MySQL/PostgreSQL (e.g. runs missing the sorted metric always sort last).Testing
integration-tests-v1and V2 api tests to utilize both databases (with cache matrices). Unit coverage expanded.Migration Guide
pipeline-install-config.data.postgresExtraParamswith ansslmodebefore applyingplatform-agnostic-postgresql. The base overlay intentionally fails closed when this is omitted; see its README for an example.NULL) are now always ordered last, in both ascending and descending order, consistently across MySQL and PostgreSQL. Previously this relied on each database's defaultNULLplacement, so MySQL placed such runs first in ascending order. Only the relative position of runs missing the sorted metric changes; runs that have the metric are unaffected. This also fixes a bug where paging past a run without the selected metric failed withcannot sort by fieldinstead of returning the next page. Sorting by regular fields (name, timestamps, etc.) is unchanged.Preceding PRs
Follow-up Issues, PRs, and Discussions
ac3a4c656generate invalid SQL after upgrade #13858 & fix(backend): decouple metric name from SQL column in list page token #13547Guide for Reviewers