Canopy - tags, labels, and reaching outside the image

Canopy's mechanism, from the first part of this pair, in short: a singleton that is itself a tree, domains hanging off it by name, and at every leaf either a raw CanopyCell or a CanopyAccessor binding that position to a live object through two independently optional selectors - one for reading, one for writing. Config and metrics turn out to be the same mechanism read from opposite ends, and two trees land on top of each other through two deliberately different operations: merge, which combines two trees structurally and refuses anything that doesn't fit cleanly, and apply, which pushes values from one tree into another wherever the receiving side already has a matching position, silently skipping the rest.

Two questions stayed open there. What happens when two sources - a config import and a component's own metrics, say - genuinely need to share one subtree, and still be told apart afterward? And once a value is honestly global inside the image, is there any reason it should stop being reachable at the image's own process boundary? Tags answer the first. The rest of this piece answers the second.

Domains are structural, tags are overlapping

Merge still needs a notion of ownership, though - two things landing in the same domain have to agree on who's responsible for which part of it, or a reload can't tell what's safe to redo. A domain already gives you one axis of that: it's the coarse, structural one, a top-level position in the tree that exists before anything gets imported into it. But ownership inside a domain doesn't stay that clean. A config file importing into #apptiveGrid and the metrics registered on the same domain both touch the same subtree, and neither "give config its own domain" nor "give metrics its own domain" works, because the whole point was that a config leaf and a metric leaf can be the very same position.

That's what tags are for: ownership within a domain, and unlike a domain, a tag is allowed to overlap with another one over the same nodes. Every CanopyNode carries a tags set, and importing a tree lets you stamp every node in it at once:

"CanopyBranchNode"
importTree: anObject tags: aCollection
    | tree |
    tree := anObject asCanopy.
    tree recursiveDo: [ :each | each addTags: aCollection ].
    self merge: tree.
    ^ tree

Import a config file tagged #config, register a component's metrics tagged #metrics, and merge both into the same domain - a key that exists in both carries both tags, a key that only one import touched carries only its own. The domain stayed one shared tree; the tags are what let you still tell the two contributions apart afterward.

A config import tagged #config and a metrics registration tagged #metrics, each contributing their own keys, merge into one #apptiveGrid domain: #logging keeps only its config tag, #requestCount keeps only its metrics tag, and #cache - present in both - carries both

That separability is the entire reason tags exist, because it lets apply: become selective: apply:withTag: walks the tree exactly the way apply: already did in the first part, except a leaf only actually gets pushed if value tags includes: aTag - one extra check, nothing else about it changes.

A config file changes on disk and gets reloaded - apply:withTag: #config pushes only the values that came from a config import back in, leaving whatever the metrics side contributed on the same nodes untouched. An export only ever wants the #metrics side, never the #config one, even though by the time either operation runs they're sitting in the exact same branch. Domain says which tree; tag says which part of it this particular operation actually owns.

A position in the tree, and what's actually there

Tags live on the node, not on whatever value happens to sit there - a tag is about who put something at that spot, not about the value itself. That distinction is worth making precise, because the next two sections only make sense once it's back in view: a leaf in Canopy's tree is really two separate objects, not one.

One of them is the position: where in the tree this is, what its key is, what's above it, which tags mark it as belonging to. The other is what's actually there right now: the value, read fresh every time something asks. Canopy keeps the two apart on purpose - in code terms, a CanopyBranchNode, CanopyCell or CanopyAccessor is the position, and a CanopyValue or CanopyMetric is the holder a read currently returns, a short-lived object built new on every access rather than something the node stores.

A CanopyBranchNode (#virtualMachine) and its child CanopyAccessor (#memorySize) connected by the tree's own parent/child edge; the accessor separately points at a CanopyMetric holder, carrying the current value and a description, re-read fresh on every access rather than stored on the node itself

Labels and, as of recently, descriptions both turn out to live on either side of that split - one half known by the position, the other by whatever the read side just answered - which is exactly why each of them needs two sources instead of one.

That wrapping buys something bigger than labels and descriptions, though. The moment every leaf answers with a holder instead of a bare value, any textual view over the tree can ask the same small set of questions of any leaf at all, whatever kind of value actually sits there - what's your value, what do you call yourself, what do you know about yourself - and get something back. A tree of bare values would have nothing to say for itself beyond the numbers; a tree of holders can explain itself, to a Prometheus scrape, to a browser, to whatever reads it next. The two sections after this one are both just that same question, asked by a different reader.

One label at the root, on every line underneath

Start with the reader that asks first, in production: a Prometheus scrape. A Prometheus line doesn't just carry a name and a number, it usually carries labels too - status="500", say - and those labels can come from two genuinely different places that happen to need the same mechanism. A metric's own labels are something the component knows about itself: a status code, a database name, whatever varies between its readings. But some labels aren't about the component at all, they're about the environment the whole image happens to be running in - which stage, which region, which deployment. Writing that second kind of label into every single metric method by hand would mean every component ends up knowing something that's none of its business.

Canopy solves it the same way it solves the name: CanopyNode>>labels walks the parent chain exactly the way allKeys already does to build a dotted path, except it concatenates label dictionaries instead of key segments:

"CanopyNode"
labels
    "The labels of this node and of every node above it - the same walk up the parent chain
    that allKeys does for the name. A label set closer to the value wins over an inherited one,
    so a single stage label at the root reaches every line without being repeated anywhere."
    ^ parent
        ifNotNil: [ parent labels , labels ]
        ifNil: [ labels ]

Setting one is just as plain: labelAt:put: stores it on that node alone, and every node below it inherits it through labels. One call at startup, on the root:

Canopy instance labelAt: #stage put: ApptiveGrid deploymentMode.

and every metric exported anywhere in the tree carries stage="production" (or whatever the image's own deployment mode happens to be) - including a domain that gets registered only later, because labels is computed by walking up at export time, not copied down once when something is added. The component that owns the actual metric never has to know #stage exists; the environment-level fact lives exactly once, at the one node that's actually responsible for knowing it.

One #stage label set once at the Canopy root reaches two completely unrelated leaves - #image/#processesActive and #virtualMachine/#memorySize - because labels walk up from wherever a leaf happens to sit, not because either branch was told about the other

The two sources meet at the exporter, which combines a holder's own labels (what the component knows) with the node's inherited ones (what the environment knows) into one set for each line - node labels , holder labels - and wherever the two disagree on a name, the holder's own entry wins, so a label set closer to the actual value always takes precedence over one inherited from further up. Output-wise, labels are written sorted by name - since which label got registered first is otherwise just insertion-order noise, not something that should ever change what a scraped line looks like - and a leaf with no labels at all is written with no braces, rather than an empty, slightly suspicious {}.

A description follows the identical pattern, with the precedence flipped. Every node - branch, cell, or accessor - and every value holder can carry one; the HELP line takes the node's own description first, falls back to the holder's, and only then to the plain key name. Labels accumulate from the holder up through the tree; a description, read down from the node, simply overrides.

Taking "global" all the way

If a value is genuinely global and reachable from anywhere in the image, there's no principled reason "anywhere" should stop at the image's own process boundary. Canopy doesn't stop there. CanopyBrowseHandler is a plain HTTP handler that resolves a URL path segment by segment through the exact same tree navigation everything else in Canopy uses, answering with whatever it finds, or a 404 if the path doesn't resolve to anything.

GET /pharo/virtualMachine/memorySize walks the tree the same three steps a Smalltalk caller would take with Canopy / #pharo / #virtualMachine / #memorySize, just spelled as URL segments instead of message sends. PUT on the same path pushes a value in through the identical apply:/value: mechanism a config-file reload uses - it just arrives over HTTP instead of from disk. On its own, mounting it is one line, JSON in and out:

server delegate: CanopyBrowseHandler new.

A browser-friendly decorator on top, CanopyBrowseHtmlHandler, renders a branch as a table and a leaf as an editable field instead, and getting the whole tree of a running image up as a page anyone on the network can read and edit still takes three lines:

Canopy registerSystemMetrics.
| server |
server := ZnServer startOn: 8080.
server delegate: CanopyBrowseHtmlHandler new.

That's what taking "global value" seriously actually buys you. A config value that used to die the moment its one loading pass finished is now a URL. A metric that used to exist only for the instant an export line got written is the same URL, read instead of written. Once a value is honestly modeled as something that lives at a stable, named position for as long as the image runs, there's nothing left to specially build to make it reachable from a debug console, a health check, or a browser on the other side of the network - it already was.

The browse handler is the general-purpose, read-and-write version of that idea; the one actually running in production is narrower on purpose. CanopyMetricsHandler answers only the metrics half - a Prometheus scrape target, GET-only, text instead of JSON:

"CanopyMetricsHandler"
prometheusText
    "With no configured domains, export the whole Canopy tree (all domains, unprefixed) - the
    default. With domains configured, export only those, each with its optional prefix, resolving
    each domain fresh so metric values stay live and re-registration is picked up."
    | visitor |
    exports isEmpty ifTrue: [ ^ self formatter format: Canopy instance ].
    visitor := self formatter.
    exports do: [ :assoc |
        assoc value
            ifNil: [ visitor addDomain: (Canopy / assoc key) ]
            ifNotNil: [ visitor addDomain: (Canopy / assoc key) prefix: assoc value ] ].
    ^ visitor export

Mounting it is as small as the browse handler was - point a ZnServer at an instance, or delegate /metrics to one on an existing one:

server delegate: CanopyMetricsHandler new.

With nothing configured it exports every domain currently registered, unprefixed - which is exactly the "everything the application currently knows about itself" facade from the first part. addDomain:/addDomain:prefix: narrow that down when a deployment only wants some of it exposed, or wants an old naming scheme kept alive at the boundary without it leaking back into the tree itself. Either way, each scrape resolves the domain fresh through the ordinary Canopy / key lookup - the same single tree, read live, every time something asks.

The honest version of the original problem was never "how do I read a config file" or "how do I export a metric." It was "I have values that are global, mostly read-only from the consumer's side, and need to live somewhere every part of the image can reach them." Config and metrics are just two directions you can walk that same sentence. Once you say it that plainly, installing one small tree that both can hang off of - taggable where ownership overlaps, labeled where the environment knows something the component doesn't, reachable from a browser just as readily as from a Smalltalk message send - stops looking like a simplification you're choosing and starts looking like the thing you should have built the first time.

The code is on github if you want to read ahead.

Found something worth flagging? Send feedback.
This work is licensed under CC BY 4.0.