Repository navigation
Support DiscoveryNotify Handling #53
Description
Activity
Not explicitly mentioned, but a bridge endpoint may send a DiscoveryNotify to the BO to notify the BO of routing table changes (such as a new endpoint getting added/removed downstream to the bridge). In reaction to the discovery notify sent by a bridge, the BO should re-query the routing table from the bridge.
What section(s) of DSP0236 do you feel substantiate this?
Not explicitly mentioned, but a bridge endpoint may send a DiscoveryNotify to the BO to notify the BO of routing table changes (such as a new endpoint getting added/removed downstream to the bridge). In reaction to the discovery notify sent by a bridge, the BO should re-query the routing table from the bridge.
What section(s) of DSP0236 do you feel substantiate this?
They don't, explicitly. This was something suggested by the editor of DSP0236 and we have bridges that implement this command to indicate endpoints downstream of the bridge falling off the network. I can ask for it to be added to the spec. itself.
I am not a 100% sure that the control daemon should directly re-query the routing table on its own accord upon receiving the DiscoveryNotify message, but do you think it is reasonable for it to emit it as a signal on the bus owner D-Bus interface?
An external entity such as the reactor could use this signal as a trigger to relearn the endpoint?
So, I think it's important to keep in mind that while a given node in the network will have enough information to route a packet to a given EID, there's no signalling mechanism for that EID being present in the network beyond its immediate bus owner. Essentially, when a BO (let's call it
A) allocates a range of EIDs to a local bridge device (B), it adds the routes to its local route table, but it has no visibility below the bridge as to whether or how they're assigned to some device below the bridge (C, say).Put another way, there's no strict global consistency for the route model of the network topology: Each node may have a different shape to its route table, based on what it knows from its perspective in the network.
There's some more exploration of this in a discussion with the Ampere folk on the OpenBMC discord
Further, MCTP is defined to drop packets with no notification to the sender on error. From DSP0236 v1.3.3, 8.7:
- Un-routable EID
An MCTP bridge receives an EID that the bridge is not able to route (for example, because the
bridge did not have a routing table entry for the given endpoint).
Together, I think it's an extra complication that's (currently) outside the spec to try to back-propagate route table changes below a bridge to the bridge's own BO, e.g. by repurposing
Discovery Notify. Rather,Aalways routes packets forCtoBbased on its route table entry that exists due to the EID pool allocation it (A) gave toB, andBdrops the packets forCif it knows that no allocation has been made forC.I think the only flow that's necessary upon receiving
Discovery Notifyis for the BO to issueEndpoint Discoveryin response (though, not necessarily excluding a priorPrepare for Endpoint Discovery, but I don't yet see that it should be necessary).- Un-routable EID
So Discovery Notify not implemented at all at the moment, right?
As I understand the spec, hot-plug devices need this command if the transport binding offers no way to handle hot-plug events.
While testing, I tried to use Discovery Notify events to signal the presence of a device on a serial interface, which didn't work.
Its not stated explicitly in the spec, but I would assume, that for serial transport, the command would be send with EID 0 for both source and destination when a hot-plug device becomes available.The spec also mentions that the command may be send as a datagram.
The behavior I observed was mctpd trying to send a reply that fails withreply_message can't take EID 0.Looks like mctpd tries to reply to the datagram, which should not happen, even in an error case right?
(I'm guessing mctpd tried to reply with a unsupported error code)mctpd: Got control request command code 13 mctpd: Ignoring unsupported command code 0x0d mctpd: BUG: reply_message can't take EID 0 mctpd: Error handling command code 0d from sockaddr_mctp_ext eid 0 net 1 type 0x00 if 4 hw len 0 0x: Protocol errorSo Discovery Notify not implemented at all at the moment, right?
That's correct; the main transport bindings we support have typically not needed it for their use-cases, but it would make sense for serial. However, there is no specific support for it mentioned in DSP0253, unlike the other transports (PCIe VDM, I3C), which do specify use of Discovery Notify. This may just be an oversight for DSP0253.
Any yes, we cannot send an EID-addressed reply to EID 0, we would need to specify physical addressing for that. This is the only instance (so far) of needing a (src = 0, dest = 0) incoming message, so we would likely need some special handling for that, and for the datagram case.
Are there any plans to support the Full Discover commands (0B and 0C) via D-Bus interfaces when MCTP is operating as a Bus Owner?
@caowg2013 : There are no transport bindings implemented, in Linux, that require the full discover commands at this point. So, no plans at the moment.
I am starting this thread to discuss how best to handle the
DiscoveryNotifymessage as defined in DSP0236 within mctpd where mctpd is the bus owner on the interface that receives the DiscoveryNotify event.Per the MCTP spec, DiscoveryNotify may be sent by an endpoint as either a request or a datagram, and is used to announce the hotplug of the device to the bus owner. For cases where the message is emnating from a non-bridge endpoint, this message may be used by the BO to perform EID assignment and discovery (Get MCTP types/UUID) on the new endpoint.
Not explicitly mentioned, but a bridge endpoint may send a DiscoveryNotify to the BO to notify the BO of routing table changes (such as a new endpoint getting added/removed downstream to the bridge). In reaction to the discovery notify sent by a bridge, the BO should re-query the routing table from the bridge.
Does the above flow make sense from the perspective of a bus owner? Is this something that can be implemented by mctpd?