Skip to content

TLS stream corruption under concurrent threaded reads on macOS arm64 #668

Description

@Torben1977

Summary

Concurrent threaded reads with Encrypt=yes can corrupt the TDS/TLS stream on macOS arm64. The issue reproduces with mssql-python 1.10.0 and also with pyodbc + Microsoft ODBC Driver 18, while the same workload is clean with Encrypt=no.

This does not appear to be a Python application pooling issue: the repro uses one private connection per thread, no SQLAlchemy pool, no shared connection/cursor objects, and synthetic constant data.

Environment

  • OS: macOS arm64, Apple Silicon M5 Pro
  • Python: 3.12
  • mssql-python: 1.10.0
  • pyodbc: 5.3.0, for comparison only
  • Microsoft ODBC Driver: msodbcsql18 18.6.2.1, for comparison only
  • unixODBC: 2.3.14, for comparison only
  • OpenSSL: 3.6.3
  • SQL Server: SQL Server 2025 RTM-CU6 in local Docker, x86_64 container under Rosetta
  • Connection: local localhost,1433, TrustServerCertificate=yes

Expected behavior

Two Python threads, each using its own independent database connection and cursor, should be able to fetch large result sets concurrently without corrupting returned row data or the protocol stream.

Actual behavior

With mssql-python and Encrypt=yes, repeated runs show intermittent failures and hangs. Examples include:

Driver Error: Communication link failure; DDBC Error: [Microsoft]The connection is no longer usable because the server response for a previously executed statement was incorrectly formatted.

In a small isolated process matrix using the same workload:

mssql-python Encrypt=yes: 4/10 clean, 4/10 corrupt, 2/10 timeout
mssql-python Encrypt=no:  10/10 clean, 0/10 corrupt, 0/10 timeout

The comparable pyodbc/msodbcsql18 repro shows the same pattern:

pyodbc + Encrypt=yes + threads: corrupt intermittently
pyodbc + Encrypt=no  + threads: clean
pyodbc + Encrypt=yes + multiprocessing: clean

Typical pyodbc comparison errors include:

[Microsoft][ODBC Driver 18 for SQL Server]Protocol error in TDS stream (0) (SQLGetData)
[Microsoft][ODBC Driver 18 for SQL Server]Unknown token received from SQL Server (0) (SQLGetData)
SSL Provider: wrong version number / record layer failure

Minimal mssql-python repro

from __future__ import annotations

import os
import sys
import threading

from mssql_python import connect

PWD = os.environ["DB_PASSWORD"]
CONN = (
    "Server=localhost,1433;"
    "Database=orglith;"
    f"UID=sa;PWD={PWD};"
    "Encrypt=yes;"
    "TrustServerCertificate=yes;"
)
Q = "SELECT TOP 200 REPLICATE(CAST('A' AS NVARCHAR(MAX)), 50000) FROM sys.all_objects a CROSS JOIN sys.all_objects b"

errors: list[tuple[str, str, str]] = []


def worker(who: str) -> None:
    try:
        with connect(CONN) as conn:
            for i in range(5):
                cur = conn.cursor()
                cur.execute(Q)
                rows = cur.fetchall()
                for n, row in enumerate(rows):
                    value = row[0]
                    if len(value) != 50000 or value.count("A") != 50000:
                        raise ValueError(f"[{who} it={i} row={n}] corrupt: len={len(value)} count={value.count('A')}")
                cur.close()
    except Exception as exc:
        errors.append((who, type(exc).__name__, str(exc)[:200]))


threads = [threading.Thread(target=worker, args=(f"t{i}",)) for i in range(2)]
for thread in threads:
    thread.start()
for thread in threads:
    thread.join()

print("RESULT:", f"CORRUPT {errors}" if errors else "CLEAN")
sys.exit(1 if errors else 0)

Run several times, or from a parent script with subprocess timeouts, because failures are intermittent.

Control observations

  • Changing only Encrypt=yes to Encrypt=no makes the workload clean in repeated runs.
  • Running the same logical workload in separate processes instead of threads is clean with TLS enabled.
  • A pure Python/OpenSSL AES-GCM stress test and a Python ssl TLS echo stress test are clean on the same machine.
  • FreeTDS/pymssql and pytds TLS comparison reads are clean against the same SQL Server.
  • Disabling pyodbc pooling does not explain the issue; the repro uses private connections and also reproduces in mssql-python.
  • MARS_Connection=Yes, PacketSize=8192, and OPENSSL_armcap=0 were not reliable fixes under larger samples.

Why I am filing this here

The issue reproduces with mssql-python 1.10.0, so it is not isolated to pyodbc or unixODBC. Because mssql-python uses DDBC and the Microsoft native SQL Server/TDS stack, this looks like a lower-level TLS/TDS stream integrity issue on this macOS arm64 path.

I have not yet confirmed whether x86_64 Linux is affected; current confirmation is macOS arm64 only.

Activity

  1. github-actions commented on Jul 8, 2026

    @github-actions

    Hi Torben1977, thank you for opening this issue!

    Our team will review it shortly. We aim to triage all new issues within 24-48 hours and get back to you.

    If you have additional information to share, please feel free to update the issue.

    Thank you for your patience!

  2. Torben1977 commented on Jul 8, 2026

    @Torben1977
    Author

    Follow-up: I was able to run the same logical repro on Azure/Linux against a real Azure SQL canary database using the same backend image that the canary tenant currently runs.

    Azure/Linux result

    Environment:

    • Azure Container Apps job
    • same image as canary tenant: orglithacr.azurecr.io/orglith-api:2026.07.5-ae80f8b
    • DB target: Azure SQL canary database
    • driver path: pyodbc + ODBC Driver 18 for SQL Server
    • connection settings: Encrypt=yes, TrustServerCertificate=no
    • workload: same 2-thread synthetic NVARCHAR(MAX) integrity check

    Result:

    run=0: CLEAN
    run=1: CLEAN
    run=2: CLEAN
    run=3: CLEAN
    run=4: CLEAN
    run=5: CLEAN
    run=6: CLEAN
    run=7: CLEAN
    run=8: CLEAN
    run=9: CLEAN
    SUMMARY: clean=10/10 corrupt=0/10 timeout=0/10 sample=[]
    

    So at least in this Azure/x86_64/Linux canary path, I could not reproduce the corruption.

    Interpretation

    This makes the current evidence much narrower:

    • confirmed broken locally on macOS arm64 under threaded TLS reads
    • not reproduced on Azure Container Apps / Linux x86_64 with the same general workload and Encrypt=yes

    So this may be a macOS/arm64-specific issue in the native Microsoft TLS/TDS path, or at least something that is not affecting the Azure/Linux x86_64 path.

    Important note

    My first Azure job run failed for a bad reason: I had passed TrustServerCertificate=false / Encrypt=true as literal strings in the repro connection string, and the Linux ODBC driver rejected that attribute format. After correcting the values to ODBC-style TrustServerCertificate=no and Encrypt=yes, the 10/10 clean run above is the valid result.

    If useful, I can also attach/inline the exact ACA job YAML used for the Linux verification.

  3. dudenamedsteve commented on Jul 8, 2026

    @dudenamedsteve

    Seeing what looks like the same underlying issue, but manifesting as partial data corruption rather than outright errors/timeouts — worth noting as a possibly related but distinct symptom of the same root cause.

    Environment:

    • OS: macOS arm64 (Apple M1)
    • Python: 3.14
    • mssql-python: 1.10.0
    • Connection: remote SQL Server, Encrypt=yes, TrustServerCertificate=yes
    • Workload: a long-running polling service with ~10-15 concurrent threads, each opening its own private connection per DB call (context-manager pattern, no shared connections/cursors, no connection pooling implemented at the application level)

    Observed behavior:
    A few times over a multi-hour session, string data returned from a SELECT (fetched via parameterized query, NVARCHAR columns) came back with a single character substituted, with the string's length otherwise unchanged and no error/exception raised on the query itself. Examples (values changed for illustration, pattern preserved exactly):

    • "INC1234567" (expected) → "IrC1234567" (actual) — 2nd character changed
    • "...last transaction timestamp..." (expected) → "...last transaction t3mestamp..." (actual) — one character changed mid-word

    Separately, and most tellingly: the same single-character-substitution signature also appeared in a pure-Python string literal that never touches the database at all — a hardcoded dict value serialized via json.dumps() and sent over a plain HTTPS POST to an unrelated third-party API. The literal was correctly spelled in source, but the runtime output showed the identical corruption pattern (single char changed, length preserved). This suggests to me that if there is a buffer-handling issue in the concurrent TLS/DDBC read path, it may be corrupting adjacent heap memory rather than only the specific buffer being read — which would explain corruption appearing in unrelated Python objects that happen to be nearby in memory, not just the string actually being fetched from SQL Server.

    I attempted to reproduce with an isolated stress test (25 threads, ~400 iterations each, simple SELECT 1 + concurrent literal-string checks) and did not trigger it in ~3 minutes — but that repro didn't fetch real string column data or run anywhere near as long as the production session where this was observed, so a negative result there isn't strong evidence either way.

    Happy to help test a fix or provide more detail if useful.

  4. self-assigned this
    on Jul 13, 2026
  5. sdebruyn commented on Jul 14, 2026

    @sdebruyn

    We are seeing what looks like the same root cause as this issue, but it surfaces as an outright SIGSEGV rather than corrupted rows or a link failure. I filed it separately as #679 with the native stack, since the symptom is different enough to be worth its own report.

    The conditions line up with what is described here: concurrent threaded reads (eight worker threads, each with its own connection, fetching in parallel), Encrypt=yes, macOS arm64, mssql-python 1.10.0. Under sustained load the process dies with no Python traceback (exit 139). The macOS crash report puts the fault in libmsodbcsql.18.dylib at SQLFetchScroll, reached through ddbc_bindings (FetchBatchData / FetchMany_wrap), with KERN_INVALID_ADDRESS at 0x0000000000000008, a dereference through a near-null pointer. Six of 29 threads were inside the driver at crash time. That matches the guess a few comments up that the concurrent read path corrupts adjacent heap memory: same corruption, but this time it lands on a pointer rather than a data byte.

    If anyone else here hits a hard crash instead of corruption, the native stack is worth capturing, because it points straight at the faulting call. On macOS the OS writes a crash report the moment the interpreter is killed by a signal:

    • Look in ~/Library/Logs/DiagnosticReports/ for python3.12-<timestamp>.ips (or <interpreter>-<timestamp>.ips).
    • The file is JSON with a quirk: the first line is metadata, everything after it is the payload. Parse the payload, take the thread whose triggered is true, and read its frames. Each frame has an imageIndex into usedImages, which maps back to the library name and offset. The exception and termination fields give the signal and fault address.
    • Running the same workload under PYTHONFAULTHANDLER=1 also prints the Python-side stack to stderr at the instant of the crash, which is quick to check without opening the report.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

triage neededFor new issues, not triaged yet.

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions