Repository navigation
perf: let a web process reuse its database connection - #946
Merged
cigamit merged 14 commits intoSep 13, 2026
Merged
Conversation
blaipr
force-pushed
the
perf/reuse-web-database-connections
branch
from
September 12, 2026 22:03
58ca233 to
eea6e6d
Compare
blaipr
force-pushed
the
perf/reuse-web-database-connections
branch
from
September 12, 2026 22:53
eea6e6d to
259b645
Compare
blaipr
force-pushed
the
perf/reuse-web-database-connections
branch
from
September 13, 2026 08:58
259b645 to
6f3f614
Compare
make ui-api-types did not reproduce the file it is documented to produce. Regenerating it changed 73,011 lines, because the committed copy carries a header banner and prettier formatting that the target never applied. The artifact and its generator had quietly diverged, which is the one thing a generated file cannot afford. generate-api-types now does what produced the committed file: openapi-typescript, then prettier, then the banner, which lives in scripts/api-types-header.txt rather than being pasted by hand. Regeneration is byte identical afterwards. With that true, CI can check it: ui-api-types-drift regenerates the types and fails if the committed copy differs, so a serializer change that nobody followed with make ui-api-types stops being invisible. The types are regenerated here, which drops the insights credential kind removed earlier in this series: the first real drift the check would have caught. awx/ui/.schema.json stops being tracked. It is the intermediate the generator feeds on and deletes, 2 MB of it, and it arrived in the repository by accident; .gitignore keeps it out.
There is no SecurityMiddleware and no SECURE_* settings, so nosniff, the referrer policy, the opener policy and frame options all depend on whichever proxy is in front. The development compose nginx sets two of them; what a given deployment sets is anyone's guess. The application sets them for itself now: SecurityMiddleware first in the list so its headers reach responses that later middleware short circuits, XFrameOptionsMiddleware last, and four settings to go with them. A deployment that also sets them at the proxy gets the same values, not a conflict. HSTS is deliberately left out. Django only emits it on a request it believes is HTTPS, and behind a TLS terminating proxy it sees HTTP unless SECURE_PROXY_SSL_HEADER is configured, so setting it here would be silently inert on exactly the deployments that need it. The proxy sets it. The test drives the two middlewares directly. The functional fixtures call views through APIRequestFactory, which never runs middleware, and the test database is built with --nomigrations, so a real request through the stack trips MigrationRanCheckMiddleware querying a table that does not exist.
The end-to-end suite drives the UI in a real browser and never looks at whether anyone using a screen reader could follow it. @axe-core/playwright scans the rendered page against WCAG 2 A and AA, and the spec runs with the rest of the suite, so no workflow change is needed. It compares what a screen violates against a recorded baseline and fails only on something new, because a zero violation gate on an application that does not pass yet would fail on arrival and be switched off within a week. No baseline file ships: all three screens scanned, login, jobs and templates, are clean. They are clean because the scan found something real and it is fixed here rather than recorded. index.html carried no title, and Login.tsx only sets one once its branding request resolves, so until then the document had no title at all: axe reports document-title, a screen reader announces nothing useful, and the browser tab shows the URL. index.html now ships a static title, which the branded one still overrides. Verified against a development stack built from this branch: the violation reproduced before the fix and all three screens pass after it.
The type checker runs advisory over the whole tree, which is right while 2,896 diagnostics are being worked through, but it means a package that reports nothing today can start reporting tomorrow and nobody notices. typecheck-strict is the ratchet: the same checker, failing, over the twelve packages that already report nothing. The advisory run still covers everything else, and a package moves up here the moment it comes clean. Where the diagnostics actually are, for whoever works the backlog next: tests 1,107, models 603, api/views 144, api/serializers 112, sso tests 92, credential_plugins 89, tasks 79.
Nothing says what is inside a release image or where it came from. A user
who pulls ghcr.io/ctrliq/ascender:25.6.2 has the registry's word for it and
nothing else.
stage.yml, after the multi-arch manifest exists: read its digest, generate
an SPDX SBOM with syft, and attest both the SBOM and the build provenance
to the registry. promote.yml signs the digest the release points at with
cosign.
Everything is keyed on the digest rather than the tag. A tag can be moved,
including by promote.yml when it points :latest at a new release, and a
signature over a tag would then describe something else.
No keys to hold. attest-* and cosign sign through Sigstore with the
workflow's own OIDC identity, which is why both jobs gain id-token: write.
Verification needs no secret either:
cosign verify ghcr.io/ctrliq/ascender@<digest> \
--certificate-identity-regexp '^https://github.com/ctrliq/ascender/' \
--certificate-oidc-issuer https://token.actions.githubusercontent.com
Every action is pinned to a commit with its version in the comment, as the
rest of the workflows are.
There is no React.lazy anywhere, so every screen, and everything it imports, shipped in the first bundle: 3.21 MB of JavaScript before the login form can render, including ace-builds for an editor most sessions never open. routeConfig declares its twenty-five screens with React.lazy, App does the same for the two it routes to itself, and both Routes blocks render inside a Suspense boundary with the existing ContentLoading as the fallback. Login stays eager, since an unauthenticated visitor needs it immediately. Measured with vite build, before and after: chunks 16 -> 130 first bundle 3.21 MB -> 381 KB, 88% smaller total JS 4.46 MB -> 4.52 MB The total grows slightly, which is the cost of chunking, and the first paint stops paying for screens nobody opened. ace-builds is now its own 684 KB chunk that loads when an editor appears. Both UI suites pass unchanged: 177 files and 1,054 tests general, 374 files and 1,937 tests screens. eslint, prettier and tsc clean.
The setting notifications and job output links are built from still carried the Tower name, in the API, in the UI settings screen and in 56 places in the tree. Renamed, with the stored value carried across by conf migration 0011, which uses the same rename_setting helper the three earlier setting renames used. A deployment that sets the old name in /etc/tower/conf.d keeps working for a release through a fallback in production.py; the installers set it through the API, which the migration covers. Also renamed, in the same sweep: - the TOWER_SECRET_KEY environment variable that regenerate_secret_key reads, now ASCENDER_SECRET_KEY, with the old name still honoured - the Ascender credential type injects ASCENDER_HOST, ASCENDER_USERNAME, ASCENDER_PASSWORD, ASCENDER_VERIFY_SSL and ASCENDER_OAUTH_TOKEN beside the TOWER_ and CONTROLLER_ names it already sets, so ascender-kit and the collection have something to move to Left alone deliberately: LOG_AGGREGATOR_TOWER_UUID and the tower_uuid field in the log payload. That field is a wire format that Ledger and any other aggregator read, so renaming it belongs with the logging work rather than here. Verified against a running deployment rather than in theory: seeded TOWER_URL_BASE in the database, ran awx-manage migrate conf, and the value arrived under the new key with the old row gone; the settings API then serves ASCENDER_URL_BASE and no longer offers the old name.
Every response embeds a copy of each related object under summary_fields, so a page of fifty job templates carries fifty copies of the same organization, inventory, project and capability map. Nothing can ask for less, because summary_fields is built unconditionally in the base serializer. A summary_fields query parameter now says what is wanted: absent means every field, exactly as before, so nothing changes for an existing caller. "none", "false" or "0" asks for an empty object, and a comma separated list asks for those keys only. The parameter joins the reserved names, since without that the filter backend reads it as a field to filter on and answers 400. Measured on a running deployment, a fifty item job template list: default 178,835 bytes 0.521 s summary_fields=none 133,158 bytes 0.415 s summary_fields=organization,user_cap.. 140,731 bytes 26% smaller and a fifth quicker, on test data whose summary fields are nearly empty. A real deployment carries recent_jobs, credentials, the last job and the user who made it, so the saving there is larger. Asking for none skips building the fields at all, which is where the time goes; asking for a subset still builds them and then filters, so it saves bytes rather than queries. Narrowing that is worth doing when something needs it.
set_statement_timeout reads PostgreSQL's statement_timeout off uwsgi's harakiri value, so a query is cancelled before the worker is killed. Off uwsgi it falls back to DATABASE_STATEMENT_TIMEOUT, which defaults to None, so no timeout is applied at all. That is fine while uwsgi is what serves. It stops being fine the moment anything else does, which is the first thing the ASGI move would change: daphne or uvicorn would serve HTTP with no statement timeout anywhere, and nothing would say so. A process that announces itself with AWX_WEB_PROCESS now gets DATABASE_STATEMENT_TIMEOUT, or 110 seconds when that is unset, which is what the uwsgi path works out to today: harakiri of 115 less a five second margin. The web supervisor programs, uwsgi and daphne both, set the variable. Task workers, management commands and migrations are untouched: they carry no marker, so they still get no timeout, because a long query there is the job rather than a symptom. Nothing changes for a uwsgi deployment: harakiri is still read first and still wins. This is step one of the path in the ASGI thread, and it is worth having whether or not the rest of it happens.
CodeEditor is the only non-test file that imports ace, and ace is the largest thing the UI ships: its chunk was 699.78 kB, 193.77 kB gzipped, the biggest in the application, because ace-builds arrives whole and its modes and themes are separate files loaded for their side effects. CodeMirror 6 ships only the pieces that are imported. The same component on the same props is now 433.07 kB, 141.01 kB gzipped, and the whole bundle drops from 4,735,923 to 4,468,204 bytes. The component's interface does not change, so no caller does: the same props, the same four modes, the same auto height between minRows and maxRows, the same Enter to edit and Escape to leave, the same debounce, the same Ctrl-F search. Two things behave better rather than the same. The editor renders only the lines in view, where ace built a line of DOM per line of content, which is what MAX_ROWS was capping. And a read-only editor is now a tab stop, because its scroll area was otherwise unreachable from the keyboard, which axe reports as scrollable-region-focusable. The colours are ace's twilight palette to the hex, but they are now CSS variables with those values as fallbacks. The four themes set the variables on .cm-editor instead of overriding rules with !important, which is what they were doing to reach ace's classes. The tests gain from it. The editor holds its document in the DOM, one element per line, so what it shows and what is typed into it are both observable: three test files stop mocking the editor away, and the assertions dropped when the suite moved to RTL, formatted JSON, yaml to json conversion, the empty-value defaults and the edit-then-submit path, come back.
The accessibility spec added with the end-to-end suite went red on the job and template lists, reporting button-name and link-name on the page header, and color-contrast on the sorted column header. The header one is real. The pending approvals badge is a PatternFly NotificationBadge, which renders a button, wrapped in a Link. The button carries an aria-hidden bell icon and, at zero pending, no text at all, so neither it nor the link around it has a name. The button also has no handler: it was a second tab stop that did nothing, since the link around it is what navigates. So the badge is the link now, through PatternFly's component prop, with an aria-label on it. One element, one tab stop, one name, and it is still an anchor, so it still opens in a new tab. The contrast one is recorded rather than fixed. The sorted column's header is drawn in the brand green, #0e8c5d, which is 4.25:1 on the table background where AA asks for 4.5. It comes from the brand colour itself, so moving it is a palette decision across four themes rather than a patch: the same green is 4.26:1 on the light themes' white, and the lighter #12a66f that fixes the dark ones is 3.13:1 there. That wants a colour picked on purpose. The baseline file the spec has always looked for now exists and carries it, with the reasoning in the spec. Checked in a browser against the development stack: the jobs list scans clean, and the templates list reports only the recorded contrast. All three accessibility specs pass.
CONN_MAX_AGE is unset, which is Django's default of 0: every request opens a PostgreSQL connection and closes it when the response is sent. The next request opens another. Nothing pools, so that is the cost of every API call before it has run a single query. Measured against the development stack, fifty requests to /api/v2/ping/ opened fifty one database sessions. Establishing one costs 5.2 ms on a container to container connection, and 5.7 ms once the application_name and statement_timeout options are in it, against 0.29 ms for a query on a connection that already exists. A process that serves HTTP now keeps its connection for sixty seconds. The same fifty requests open five sessions, one per uwsgi worker, which is a 90% cut in connection churn for no change anyone can see. Sixty seconds is deliberately short. An idle worker holds its connection until its next request notices the age, so what is held at rest is one per web worker, and a short life keeps a connection from outliving a network path that has quietly gone away. Task workers, management commands and migrations are left alone. They hold one connection for the life of a long process, or want it closed the moment the command ends, and an age limit helps neither. DATABASE_CONN_MAX_AGE overrides all of it, including for a task process, so a deployment can turn reuse off with 0 or hold connections for longer. Web is detected the way the statement timeout detects it: uwsgi answers for every deployment today, and AWX_WEB_PROCESS answers for anything that is not uwsgi, which is what will matter if the server ever becomes daphne.
blaipr
force-pushed
the
perf/reuse-web-database-connections
branch
from
September 13, 2026 14:58
6ea8b14 to
a3c28e4
Compare
…abase-connections # Conflicts: # awx/ui/e2e/accessibility-baseline.json # awx/ui/e2e/tests/accessibility.spec.js
cigamit
previously approved these changes
Sep 13, 2026
…se-connections # Conflicts: # awx/settings/production.py
Contributor
Author
|
Conflict resolved, this is mergeable again. |
cigamit
approved these changes
Sep 13, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Sits on top of the stack ending at #945. From the roadmap list in To Do: "No pooling on the database:
ATOMIC_REQUESTSis on,CONN_MAX_AGEis unset and psycopg's pool is off."The measurement first
CONN_MAX_AGEis unset, which is Django's default of0: every request opens a PostgreSQL connection and closes it when the response is sent. Nothing pools, so that is the cost of every API call before it has run a single query./api/v2/ping/, todayapplication_nameandstatement_timeoutoptions in itBoth numbers came off the development stack and its PostgreSQL, the session count from
pg_stat_database.sessionseither side of the fifty requests.The change
A process that serves HTTP keeps its connection for sixty seconds. The same fifty requests then open 5 sessions, one per uwsgi worker: a 90% cut in connection churn, for no change anyone can see.
Sixty seconds is deliberately short. An idle worker holds its connection until its next request notices the age, so what is held at rest is one per web worker, and a short life keeps a connection from outliving a network path that has quietly gone away.
Task workers, management commands and migrations are left alone: they hold one connection for the life of a long process, or want it closed the moment the command ends, and an age limit helps neither.
DATABASE_CONN_MAX_AGEoverrides all of it, including for a task process, so a deployment can turn reuse off with0or hold connections for longer.Web is detected the way #942 detects it:
uwsgianswers for every deployment today, andAWX_WEB_PROCESSanswers for anything that is not uwsgi, which is what will matter if the server ever becomes daphne.Why not psycopg's pool
A pool helps a process handling several requests at once. Under uwsgi's prefork model each worker handles one at a time, so a pool per worker would hold the same single connection with an extra dependency,
psycopg_pool, to do it. The pool becomes the right answer when the server is async, which is the Path to ASGI thread, not this.Checks
awx/main/tests/unit/settings: 49 passing, 11 of them new, covering the marker, uwsgi, the setting winning over the default,0turning reuse off, a task process being left alone, and the emptyDATABASEScases.CONN_MAX_AGE=0, web gets60alongside the statement timeout.ruff checkandruff format --checkclean.