Skip to content

ncurses: letters with diacritics are misinterpreted as function keys (ĉ → F1, ĝ → Shift-F9, č → F5, š → Shift-Tab) #2310

Description

@dnsl48

Summary

In the ncurses frontend, typing certain non-ASCII letters does not insert the character. Instead Lem processes it as an unrelated special key. Esperanto, Czech, Polish, Romanian, Hungarian and Turkish text are all affected.

The cause is a numeric collision: Lem looks up decoded Unicode code points in a table keyed by raw ncurses KEY_* constants. Those constants occupy 258–567, which is exactly the Latin Extended-A block.

Steps to reproduce

  1. Build and run the ncurses frontend (make ncurses) in any UTF-8 terminal.
  2. In a buffer, type ĉ (U+0109).

Expected: ĉ is inserted.
Actual: Lem treats the keystroke as F1 — it runs whatever F1 is bound to, or reports Key not found: F1.

More examples:

type Lem sees because
ĉ U+0109 = 265 F1 KEY_F(1) = #o411 = 265
Ĉ U+0108 = 264 F0 KEY_F(0) = #o410 = 264
ĝ U+011D = 285 Shift-F9 KEY_SF(9) = #o435 = 285
č U+010D = 269 F5 KEY_F(5) = #o415 = 269
š U+0161 = 353 Shift-Tab KEY_BTAB = #o541 = 353
ą U+0105 = 261 Right arrow KEY_RIGHT = #o405 = 261

Root cause

frontends/ncurses/key.lisp:21:

(defun char-to-key (char)
  (or (gethash (char-code char) *keycode-table*)
      (make-key :sym (string char))))

*keycode-table* (key.lisp:9) is keyed on raw values returned by wgetch — ASCII 0–127 plus ncurses KEY_* constants 258–567.

frontends/ncurses/input.lisp:34, get-key, assembles a multi-byte UTF-8 sequence into a character and then passes that decoded character to char-to-key. For any decoded code point in 258–567, the table lookup succeeds and returns a function key instead of the letter.

The two namespaces are unrelated and must not share a lookup table.

Full list of affected characters

59 letters collide (every table entry in 258–567 that is a printable letter):

Ă ă Ą ą Ć ć Ĉ ĉ Ċ ċ Č č Ď ď Đ đ Ē ē Ĕ ĕ Ė ė Ę ę Ě ě Ĝ ĝ Ğ ğ Ġ
Ŋ Ő ő Œ œ š Ũ ſ Ƃ Ƈ Ɖ ƌ Ǝ ƒ ȇ Ȍ ȍ Ȏ Ƞ ȡ Ȣ ȯ Ȱ ȱ ȴ ȵ ȶ ȷ

Proposed fix

get-key already knows whether it decoded a multi-byte sequence. A multi-byte sequence is always a character and never an ncurses keycode, so that path should bypass *keycode-table* entirely and build the key directly with make-key :sym. Single-byte input keeps using char-to-key, so all real function keys continue to work.

This is safe because utf8-bytes only reports >1 for lead bytes 0xC2–0xF4 (194–244), and the table contains no entries in that range — so no legitimate keycode lookup is lost.

Also affects the pdcurses (Windows) frontend

frontends/pdcurses/ncurses-pdcurseswin32.lisp:557 ends in (char-to-key (code-char code)) and has the same collision. It needs a separate fix, since PDCurses returns wide characters and KEY_* codes through the same channel, making the values genuinely ambiguous there.

Related

Previously reported as #656 (closed as stale, not fixed). That report reached a different conclusion — it blamed (setf (key-to-character (make-key :shift t :sym "Tab")) (code-char 353)) in src/key.lisp — but that table maps key → character and is never consulted on input. The actual culprit is *keycode-table* in the ncurses frontend.

Environment

  • Lem main @ f5ba7d8
  • ncurses frontend, SBCL, any UTF-8 terminal

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions