Introduce a representation between language and actuation
“I want to relax” does not specify a device, address or value. It expresses an activity whose appropriate environment depends on the person, room, time and other occupants. Directly mapping the phrase to a fixed command list discards those distinctions.
The proposed engine separates behaviour recognition from environmental composition. A language model can interpret and explain a request; a structured layer resolves the resulting goal against known context and capabilities. This gives engineering a place to express constraints that should not depend on how a sentence happens to be phrased.
Model behaviour together with its phase
Preparing for sleep, being asleep and waking are phases with different requirements even though they belong to one broad activity. The behaviour model therefore needs more than a class label. It must relate the activity to its stage, triggers and relevant environment template.
The source proposal includes reading, movie watching, relaxing, sleep preparation, leaving and returning home as initial template directions. These are product abstractions to be tested with users. They are not a universal behavioural taxonomy or a diagnosis of an occupant’s internal state.
Represent goals in environmental dimensions
The environment layer describes desired conditions such as brightness, glare, thermal comfort, privacy and atmosphere. Targets may be ranges rather than one exact device setting. Several actuator combinations can serve the same intention, and the available combination differs from home to home.
A reading context might require suitable task illumination without unnecessary glare. The composer would reason over available lighting and shading capabilities rather than assume every dwelling contains the same devices. Illustrative target values must remain scenario-dependent; the architecture does not establish a universal comfort prescription.
| Module | Required context | Output responsibility |
|---|---|---|
| House semantic graph | Rooms, regions, devices and relationships | A scoped account of where an activity occurs. |
| Behaviour and templates | Activity, phase and initial preferences | Environmental goals that can be explained. |
| Device capability model | Controllable properties, feedback and constraints | The feasible action space. |
| Scene composer | Goals, current context and permissions | A bounded, inspectable proposal. |
| Learning loop | Snapshots, corrections and relevant observations | Contextual preference updates for evaluation. |
Make capability and feedback part of composition
A device capability model needs to describe what can be changed, which values are valid, how feedback is obtained and what constraints apply. A command-only output and an independently reporting device offer different evidence of success. Treating both as equally observable would make the plan appear more certain than the installation allows.
The proposed composer selects an inspectable set of actions within those boundaries. Permissions, device constraints and the consequence of a proposed change determine whether an action can proceed, should be confirmed or must be deferred. A convincing explanation cannot replace an unavailable capability or missing authority.
Solve the cold start through engineering context
A new home has little behavioural history but is not an information vacuum. Design and handover already describe rooms, device locations, capabilities and reviewed scenes. Onboarding adds explicit preferences; early corrections add local evidence about what a resident actually wants.
The proposed sequence is therefore design context, handover snapshots, onboarding and a first period of correction. This is a practical alternative to requiring a large personal interaction history before the system can suggest anything useful. It also exposes which assumptions came from templates and which came from a person.
Learn preferences rather than replaying commands
A resident increasing one light after a scene could be correcting insufficient task illumination, a temporary occupancy condition or a poorly chosen scene. Simply saving the final command sequence loses the distinction. The learning loop should preserve the context and the environmental dimension being corrected.
The source architecture stages this ambition: establish the semantic graph, capabilities and templates first; add explainable composition next; then investigate learning from corrections. The proposed feedback loop is not evidence that a particular learning algorithm has already improved outcomes or that an end-to-end autonomous system is deployed.
Evaluate usefulness across the entire chain
An evaluation programme should distinguish intent interpretation, context completeness, capability feasibility, execution outcome and resident correction. A semantically correct interpretation can still produce an impossible plan; a technically successful command can still create the wrong environment.
Useful scenarios include ambiguous room references, competing occupants, missing state feedback, unsupported capabilities and interrupted connectivity. The objective is a system that makes uncertainty actionable: ask, narrow, defer or offer an alternative when the evidence does not justify the original proposal. These are proposed evaluation dimensions, not reported benchmark results.
