Repository navigation
TLS stream corruption under concurrent threaded reads on macOS arm64 #668
Description
Activity
Hi Torben1977, thank you for opening this issue!
Our team will review it shortly. We aim to triage all new issues within 24-48 hours and get back to you.
If you have additional information to share, please feel free to update the issue.
Thank you for your patience!
- addedtriage neededFor new issues, not triaged yet.For new issues, not triaged yet.
on Jul 8, 2026 Follow-up: I was able to run the same logical repro on Azure/Linux against a real Azure SQL canary database using the same backend image that the canary tenant currently runs.
Azure/Linux result
Environment:
- Azure Container Apps job
- same image as canary tenant:
orglithacr.azurecr.io/orglith-api:2026.07.5-ae80f8b - DB target: Azure SQL canary database
- driver path:
pyodbc+ODBC Driver 18 for SQL Server - connection settings:
Encrypt=yes,TrustServerCertificate=no - workload: same 2-thread synthetic
NVARCHAR(MAX)integrity check
Result:
run=0: CLEAN run=1: CLEAN run=2: CLEAN run=3: CLEAN run=4: CLEAN run=5: CLEAN run=6: CLEAN run=7: CLEAN run=8: CLEAN run=9: CLEAN SUMMARY: clean=10/10 corrupt=0/10 timeout=0/10 sample=[]So at least in this Azure/x86_64/Linux canary path, I could not reproduce the corruption.
Interpretation
This makes the current evidence much narrower:
- confirmed broken locally on macOS arm64 under threaded TLS reads
- not reproduced on Azure Container Apps / Linux x86_64 with the same general workload and
Encrypt=yes
So this may be a macOS/arm64-specific issue in the native Microsoft TLS/TDS path, or at least something that is not affecting the Azure/Linux x86_64 path.
Important note
My first Azure job run failed for a bad reason: I had passed
TrustServerCertificate=false/Encrypt=trueas literal strings in the repro connection string, and the Linux ODBC driver rejected that attribute format. After correcting the values to ODBC-styleTrustServerCertificate=noandEncrypt=yes, the 10/10 clean run above is the valid result.If useful, I can also attach/inline the exact ACA job YAML used for the Linux verification.
Seeing what looks like the same underlying issue, but manifesting as partial data corruption rather than outright errors/timeouts — worth noting as a possibly related but distinct symptom of the same root cause.
Environment:
- OS: macOS arm64 (Apple M1)
- Python: 3.14
- mssql-python: 1.10.0
- Connection: remote SQL Server, Encrypt=yes, TrustServerCertificate=yes
- Workload: a long-running polling service with ~10-15 concurrent threads, each opening its own private connection per DB call (context-manager pattern, no shared connections/cursors, no connection pooling implemented at the application level)
Observed behavior:
A few times over a multi-hour session, string data returned from a SELECT (fetched via parameterized query, NVARCHAR columns) came back with a single character substituted, with the string's length otherwise unchanged and no error/exception raised on the query itself. Examples (values changed for illustration, pattern preserved exactly):- "INC1234567" (expected) → "IrC1234567" (actual) — 2nd character changed
- "...last transaction timestamp..." (expected) → "...last transaction t3mestamp..." (actual) — one character changed mid-word
Separately, and most tellingly: the same single-character-substitution signature also appeared in a pure-Python string literal that never touches the database at all — a hardcoded dict value serialized via json.dumps() and sent over a plain HTTPS POST to an unrelated third-party API. The literal was correctly spelled in source, but the runtime output showed the identical corruption pattern (single char changed, length preserved). This suggests to me that if there is a buffer-handling issue in the concurrent TLS/DDBC read path, it may be corrupting adjacent heap memory rather than only the specific buffer being read — which would explain corruption appearing in unrelated Python objects that happen to be nearby in memory, not just the string actually being fetched from SQL Server.
I attempted to reproduce with an isolated stress test (25 threads, ~400 iterations each, simple SELECT 1 + concurrent literal-string checks) and did not trigger it in ~3 minutes — but that repro didn't fetch real string column data or run anywhere near as long as the production session where this was observed, so a negative result there isn't strong evidence either way.
Happy to help test a fix or provide more detail if useful.
We are seeing what looks like the same root cause as this issue, but it surfaces as an outright SIGSEGV rather than corrupted rows or a link failure. I filed it separately as #679 with the native stack, since the symptom is different enough to be worth its own report.
The conditions line up with what is described here: concurrent threaded reads (eight worker threads, each with its own connection, fetching in parallel),
Encrypt=yes, macOS arm64, mssql-python 1.10.0. Under sustained load the process dies with no Python traceback (exit 139). The macOS crash report puts the fault inlibmsodbcsql.18.dylibatSQLFetchScroll, reached throughddbc_bindings(FetchBatchData/FetchMany_wrap), withKERN_INVALID_ADDRESS at 0x0000000000000008, a dereference through a near-null pointer. Six of 29 threads were inside the driver at crash time. That matches the guess a few comments up that the concurrent read path corrupts adjacent heap memory: same corruption, but this time it lands on a pointer rather than a data byte.If anyone else here hits a hard crash instead of corruption, the native stack is worth capturing, because it points straight at the faulting call. On macOS the OS writes a crash report the moment the interpreter is killed by a signal:
- Look in
~/Library/Logs/DiagnosticReports/forpython3.12-<timestamp>.ips(or<interpreter>-<timestamp>.ips). - The file is JSON with a quirk: the first line is metadata, everything after it is the payload. Parse the payload, take the thread whose
triggeredistrue, and read itsframes. Each frame has animageIndexintousedImages, which maps back to the library name and offset. Theexceptionandterminationfields give the signal and fault address. - Running the same workload under
PYTHONFAULTHANDLER=1also prints the Python-side stack to stderr at the instant of the crash, which is quick to check without opening the report.
- Look in
Summary
Concurrent threaded reads with
Encrypt=yescan corrupt the TDS/TLS stream on macOS arm64. The issue reproduces withmssql-python1.10.0 and also withpyodbc+ Microsoft ODBC Driver 18, while the same workload is clean withEncrypt=no.This does not appear to be a Python application pooling issue: the repro uses one private connection per thread, no SQLAlchemy pool, no shared connection/cursor objects, and synthetic constant data.
Environment
localhost,1433,TrustServerCertificate=yesExpected behavior
Two Python threads, each using its own independent database connection and cursor, should be able to fetch large result sets concurrently without corrupting returned row data or the protocol stream.
Actual behavior
With
mssql-pythonandEncrypt=yes, repeated runs show intermittent failures and hangs. Examples include:In a small isolated process matrix using the same workload:
The comparable
pyodbc/msodbcsql18 repro shows the same pattern:Typical pyodbc comparison errors include:
Minimal mssql-python repro
Run several times, or from a parent script with subprocess timeouts, because failures are intermittent.
Control observations
Encrypt=yestoEncrypt=nomakes the workload clean in repeated runs.sslTLS echo stress test are clean on the same machine.mssql-python.MARS_Connection=Yes,PacketSize=8192, andOPENSSL_armcap=0were not reliable fixes under larger samples.Why I am filing this here
The issue reproduces with
mssql-python1.10.0, so it is not isolated to pyodbc or unixODBC. Becausemssql-pythonuses DDBC and the Microsoft native SQL Server/TDS stack, this looks like a lower-level TLS/TDS stream integrity issue on this macOS arm64 path.I have not yet confirmed whether x86_64 Linux is affected; current confirmation is macOS arm64 only.