Summary
In the ncurses frontend, typing certain non-ASCII letters does not insert the character. Instead Lem processes it as an unrelated special key. Esperanto, Czech, Polish, Romanian, Hungarian and Turkish text are all affected.
The cause is a numeric collision: Lem looks up decoded Unicode code points in a table keyed by raw ncurses KEY_* constants. Those constants occupy 258–567, which is exactly the Latin Extended-A block.
Steps to reproduce
- Build and run the ncurses frontend (
make ncurses) in any UTF-8 terminal.
- In a buffer, type
ĉ (U+0109).
Expected: ĉ is inserted.
Actual: Lem treats the keystroke as F1 — it runs whatever F1 is bound to, or reports Key not found: F1.
More examples:
| type |
Lem sees |
because |
ĉ U+0109 = 265 |
F1 |
KEY_F(1) = #o411 = 265 |
Ĉ U+0108 = 264 |
F0 |
KEY_F(0) = #o410 = 264 |
ĝ U+011D = 285 |
Shift-F9 |
KEY_SF(9) = #o435 = 285 |
č U+010D = 269 |
F5 |
KEY_F(5) = #o415 = 269 |
š U+0161 = 353 |
Shift-Tab |
KEY_BTAB = #o541 = 353 |
ą U+0105 = 261 |
Right arrow |
KEY_RIGHT = #o405 = 261 |
Root cause
frontends/ncurses/key.lisp:21:
(defun char-to-key (char)
(or (gethash (char-code char) *keycode-table*)
(make-key :sym (string char))))
*keycode-table* (key.lisp:9) is keyed on raw values returned by wgetch — ASCII 0–127 plus ncurses KEY_* constants 258–567.
frontends/ncurses/input.lisp:34, get-key, assembles a multi-byte UTF-8 sequence into a character and then passes that decoded character to char-to-key. For any decoded code point in 258–567, the table lookup succeeds and returns a function key instead of the letter.
The two namespaces are unrelated and must not share a lookup table.
Full list of affected characters
59 letters collide (every table entry in 258–567 that is a printable letter):
Ă ă Ą ą Ć ć Ĉ ĉ Ċ ċ Č č Ď ď Đ đ Ē ē Ĕ ĕ Ė ė Ę ę Ě ě Ĝ ĝ Ğ ğ Ġ
Ŋ Ő ő Œ œ š Ũ ſ Ƃ Ƈ Ɖ ƌ Ǝ ƒ ȇ Ȍ ȍ Ȏ Ƞ ȡ Ȣ ȯ Ȱ ȱ ȴ ȵ ȶ ȷ
Proposed fix
get-key already knows whether it decoded a multi-byte sequence. A multi-byte sequence is always a character and never an ncurses keycode, so that path should bypass *keycode-table* entirely and build the key directly with make-key :sym. Single-byte input keeps using char-to-key, so all real function keys continue to work.
This is safe because utf8-bytes only reports >1 for lead bytes 0xC2–0xF4 (194–244), and the table contains no entries in that range — so no legitimate keycode lookup is lost.
Also affects the pdcurses (Windows) frontend
frontends/pdcurses/ncurses-pdcurseswin32.lisp:557 ends in (char-to-key (code-char code)) and has the same collision. It needs a separate fix, since PDCurses returns wide characters and KEY_* codes through the same channel, making the values genuinely ambiguous there.
Related
Previously reported as #656 (closed as stale, not fixed). That report reached a different conclusion — it blamed (setf (key-to-character (make-key :shift t :sym "Tab")) (code-char 353)) in src/key.lisp — but that table maps key → character and is never consulted on input. The actual culprit is *keycode-table* in the ncurses frontend.
Environment
- Lem
main @ f5ba7d8
- ncurses frontend, SBCL, any UTF-8 terminal
Summary
In the ncurses frontend, typing certain non-ASCII letters does not insert the character. Instead Lem processes it as an unrelated special key. Esperanto, Czech, Polish, Romanian, Hungarian and Turkish text are all affected.
The cause is a numeric collision: Lem looks up decoded Unicode code points in a table keyed by raw ncurses
KEY_*constants. Those constants occupy 258–567, which is exactly the Latin Extended-A block.Steps to reproduce
make ncurses) in any UTF-8 terminal.ĉ(U+0109).Expected:
ĉis inserted.Actual: Lem treats the keystroke as F1 — it runs whatever F1 is bound to, or reports
Key not found: F1.More examples:
ĉU+0109 = 265KEY_F(1)=#o411= 265ĈU+0108 = 264KEY_F(0)=#o410= 264ĝU+011D = 285KEY_SF(9)=#o435= 285čU+010D = 269KEY_F(5)=#o415= 269šU+0161 = 353KEY_BTAB=#o541= 353ąU+0105 = 261KEY_RIGHT=#o405= 261Root cause
frontends/ncurses/key.lisp:21:*keycode-table*(key.lisp:9) is keyed on raw values returned bywgetch— ASCII 0–127 plus ncursesKEY_*constants 258–567.frontends/ncurses/input.lisp:34,get-key, assembles a multi-byte UTF-8 sequence into a character and then passes that decoded character tochar-to-key. For any decoded code point in 258–567, the table lookup succeeds and returns a function key instead of the letter.The two namespaces are unrelated and must not share a lookup table.
Full list of affected characters
59 letters collide (every table entry in 258–567 that is a printable letter):
Proposed fix
get-keyalready knows whether it decoded a multi-byte sequence. A multi-byte sequence is always a character and never an ncurses keycode, so that path should bypass*keycode-table*entirely and build the key directly withmake-key :sym. Single-byte input keeps usingchar-to-key, so all real function keys continue to work.This is safe because
utf8-bytesonly reports >1 for lead bytes0xC2–0xF4(194–244), and the table contains no entries in that range — so no legitimate keycode lookup is lost.Also affects the pdcurses (Windows) frontend
frontends/pdcurses/ncurses-pdcurseswin32.lisp:557ends in(char-to-key (code-char code))and has the same collision. It needs a separate fix, since PDCurses returns wide characters andKEY_*codes through the same channel, making the values genuinely ambiguous there.Related
Previously reported as #656 (closed as stale, not fixed). That report reached a different conclusion — it blamed
(setf (key-to-character (make-key :shift t :sym "Tab")) (code-char 353))insrc/key.lisp— but that table maps key → character and is never consulted on input. The actual culprit is*keycode-table*in the ncurses frontend.Environment
main@ f5ba7d8