The last five years of AI progress have been a story about digital data, text, code, and images, processed in datacentres with huge amounts of computational power, where latency has rarely been the binding constraint. But for real-world applications in physical space, data doesn't arrive as clean paragraphs or well-lit photos. It arrives as a flood of asynchronous, noisy, physically grounded signals: an accelerometer sampling at 400Hz next to a temperature sensor updating every 3 seconds, with camera feeds arriving at 30–120Hz. This is the reality of the physical world, and today's foundation models were not built for it.
This essay looks to define Sensor Foundational Models (SFMs): a distinct class of foundation model built to reason and act, in real-time, over asynchronous and heterogeneous sensor streams, under the power and compute constraints of edge hardware.
I want to start with what SFMs are not. They are not a scaled-down version of a language model, shrunk to run on low-compute hardware.
Instead, SFMs solve a different problem: fusing signals that disagree on timing and modality, producing decisions in real-time (tens of milliseconds), and doing it, ideally, on hardware that's battery powered rather than tethered to the grid or a diesel generator. I believe that SFMs are a layer the current AI stack is missing for bringing AI into the physical world, and that the underlying technology is finally ready. I will also use this essay to talk about where I think the category gets built and won, from the home, to drones, to the factory floor, to space.
It is important to recognise that there are transformer architectures that are specifically designed for asynchronous time series data such as ContiFormer1 or even models designed for multimodal fusion or event-based processing. However, I don’t believe there is any dominant foundational model that brings together all the requirements I think are needed for bringing AI truly into the physical world.
The Physical World Is AI's Next Frontier
Every major foundation model to date has been trained on a record of the world, not the world itself. Text is a compressed record of human thought. Images are a static record of a scene. Even video, the most "physical" of the standard modalities, is a dense, synchronous, evenly sampled record. It's a format chosen because cameras happen to produce it, not because it's how the physical world presents itself to a machine. (As it happens, video is a key component of some SFMs.)
Sensors, combined together, don't work that way. A modern embedded system such as a drone contains multiple sensors that each sample a different physical quantity, at a different rate, with a different noise profile, and that go silent or lag unpredictably depending on power state, radio contention, or physical occlusion. An IMU might produce a reading every 2 milliseconds. A CO2 sensor might produce one every 3 seconds. None of this looks like a token stream, and none of it waits politely for the others before the system has to make a decision.
This asynchronous, heterogeneous, real-time character is precisely what today's foundation models aren't designed to handle. Mainstream LLMs and vision transformers typically assume tokenized, ordered, or frame-based inputs, though transformer variants are increasingly being proposed for irregular multimodal fusion. Both assume you can afford to wait for the whole input before responding, and both assume you have a GPU cluster, not a lithium-ion battery, to do the responding with. The result is a large and growing gap: the physical world generates more raw signals every second than the digital world ever will, and almost none of the current foundation model stack is built to live inside it.
I believe we're limiting our devices' ability to "reason" by not giving them enough "senses." Take a drone as an example. The main way to control one today is via input from a human monitoring a video feed. I don't think that's how it should be flying. A drone should have an instructed mission, and should fuse sensor data together, from an IMU, multiple video streams, pressure sensors, accelerometers, and more. That's what we do as humans, with eyes, ears, nose, and sensation on our skin, albeit with analogue data.
Defining SFM
Sensor Foundational Model (SFM): a foundation model trained to perform real-time reasoning and actioning over asynchronous, heterogeneous sensor data streams, under the power, memory, and latency constraints of edge hardware.
Another, longer way of articulating this is that an SFM implementation should be designed and tested against explicit targets such as sub-500ms response latency, single-digit-watt sustained power draw, a memory footprint found in consumer electronics (sub 12GB), quantization on commodity edge silicon, and performance that holds up even when a sensor input drops out or arrives late.
Each of the following terms are crucial as they rule out common but insufficient alternatives.
Real-time. Not "fast inference" in the sense of a low-latency cloud API call. Real-time means the model's operating budget is set by the physical process it's embedded in. For example, a drone's control loop, a factory safety interlock or a fall-detection algorithm, not by what feels responsive to a person waiting for a chat reply. Missing the required action by being slow isn't a just bad user experience, it could have knock-on consequences. For feasibility's sake, I'd define real-time as anything under 500ms, with a target of sub-100ms.
Reasoning and actioning. Most sensor machine learning to date stops at classification or forecasting, eg, is this a fall, what is the likelihood this machine will fail in the next week, what will tomorrow's energy price be. An SFM should close the loop by turning an inference into a decision via reasoning and then an action, under uncertainty and incomplete information.
Asynchronous, heterogeneous streams. Sensor data comes in different modalities at different sampling rates. SFMs are designed to take that in and process it in real-time. This is what I think will help bring AI out of the computer screen and into the real world.
Edge hardware constraints. Power, memory, and thermal budgets shouldn’t be an afterthought to be solved by compressing a cloud model after the fact. Power and memory should be first-class design constraints. A model that reasons brilliantly but needs a server rack to do so hasn't solved what SFMs are meant to solve.
Put together, these four properties describe a different problem from the one that produced GPT, Gemini, or Claude. The problem happens to share the "train a large general-purpose model, then adapt it to many tasks" idea that made foundation models work in the first place. That shared adaptation is why the term foundation model still applies, and why I think it's worth keeping SFMs in the same conceptual family rather than treating them as a niche subfield of embedded ML.
Why General-Purpose Models Don't Just Transfer
I want to anticipate a reasonable objection: hasn't this already been solved by distilling large models onto power and compute-constrained hardware? I don’t believe so. There are two main reasons no amount of distillation can fix, because they're about the type of problem, not the size of the model.
Time-series foundation models such as Chronos, MOIRAI, and Lag-Llama are capable at forecasting, but forecasting a relatively clean, single-modality stream is a narrower and easier problem than fusing irregular, multimodal with possibly partially-missing sensor data into a real-time decision. Large multimodal models remain optimised for structured vision-language pairs, not open-ended sensor fusion. Many TinyML deployments remain task-specific and require substantial per-device or per-environment tuning. A recent academic study2 benchmarking time-series foundation models on building-sensor data found that generalisation across different buildings and deployment conditions remained incomplete, and that naively combining multiple sensor modalities could hurt performance rather than help it. No existing approach yet handles cross-context generalisation from irregular, asynchronous, multimodal sensor data under edge power constraints, in a single architecture. That's the gap I am naming with SFMs, not inventing.3,4,5,6
Transformer architectures, and the tokenisation schemes built on top of them, generally assume a well-ordered sequence. True sensor fusion requires reasoning over streams that are out of sync with each other. That asynchrony carries information, and often gets thrown away the moment you force everything onto a shared clock. This is because the majority of current architectures make asynchronous data conform to the model’s representation rather than treating the asynchronous observation as a critical part of its understanding of the real world, and utilising the ‘mess’ of data to its advantage.
For models working off sensor date in “real-time” this is important, because a hallucinating SFM controlling a drone or a medical wearable produces a wrong action in the physical world, often with no chance to catch the error before it has consequences. That raises the bar for graceful degradation when a sensor drops out. None of these are solved problems by today's general-purpose foundation models, which were built for a domain where a wrong answer is embarrassing rather than dangerous (I will note that there are AI models being used in life/death situations by the military, but it is my understanding that humans are still, as of writing, kept in the loop).
The closest existing research to what I’m describing is JEPA, (Joint Embedding Predictive Architecture), the world-model approach Meta’s research group has been building out. A JEPA-style model predicts the next state of the world in a compressed, learned representation space, then checks that prediction against what actually happens. Meta’s V-JEPA 2, released this year, showed this approach can support understanding, prediction, and planning from video alone.7 It wasn’t built for asynchronous, multi-rate sensor fusion however the core idea still applies of learning to predict what happens next in an abstract representation of the world which is close to what an SFM needs to do with an different heterogeneous multimodal sensors.
The area of work that I would like to raise is VLA’s (Vision Language Action models) which have become a dominant approach in robotics over the past couple of years.8 A VLA takes in an image and a language instruction and outputs a robotic action directly which is the same reasoning/action loop SFMs would close. Where VLAs and SFMs differ is in scope and constraint. Most VLA work is still built on two modalities, vision and language, and runs with large compute budgets rather than edge or consumer grade hardware. An SFM should handle a multiple heterogeneous streams on ideally a sub-10-watt budget. VLAs prove the action half of the problem is tractable. Extending that closed loop to arbitrary sensor types, running on limited power, is a new challenge.
As an example of what an SFM might actually look like under the hood, I’d start with a separate, small encoder for each sensor type such as one that learns to turn a raw IMU signal into a shared representation, another for radar, and another for a slow-changing value like temperature, all feeding into a JEPA-style core that predicts the next state of that shared representation as new readings arrive. What that predicted state does next, how it turns into a control output for a drone or an alert on a wearable is where it will get very interesting, and is for now open. Bolting on a specific action mechanism this early would tie the definition to one deployment before the category has had a chance to be built by more than one team. I think that’s a research problem in its own right, one that gets solved collectively as more people start building toward this category over the next few years.
The Four Pillars of an SFM
To build upon the definition I proposed earlier in this essay, I believe any credible SFM architecture has to make deliberate choices in four areas that general-purpose foundation models get to take for granted.
Multimodal native. This is what will make SFMs the gold standard in the physical world. Multimodal functionality lets an SFM take in multiple sources of input, analyse and understand them, and then act on what it's observed. Without multimodal functionality, SFMs would fit into existing classes of unimodal models, like Chronos or MOIRAI.
Asynchronous data fusion. Having multimodal sensors inherently means the data is asynchronous. These streams need to work harmoniously together, and asynchronous data fusion is core to making that happen. The sampling rate of a radar at 10MHz, a camera at 30Hz, and an IMU at 32kHz all need to work together in harmony, even if one drops out. Part of this pillar is building in redundancy, so the SFM doesn't catastrophically fail if there's a lapse in one of the modalities of data.
Real-time analysis and action. At their core, SFMs should be designed to have as close to real-time analysis as possible, which can then be acted on. This works in conjunction with the pillar above: the action produces a real-world impact, and the sensor data feeds back into the model, either reinforcing the action or correcting it, helping the SFM adjust in real-time.
Local first. To be truly useful, SFMs need to run locally. In the short term, high-speed data links will still be necessary while low-powered hardware races to keep up and the models get optimised, but the north star is locally run SFMs making fully autonomous decisions. This is imperative for the future of some industries, especially defence, where the RF spectrum will be compromised and communication can't be guaranteed. Practically, that means the hardware should be battery operable and of reasonable compute. I would suggest sub 12GB RAM and sub 10W sustained power draw. This of course has nuance to it. There will be deployments that require more computation, and thus require more demanding hardware capabilities. In these scenarios, power might not be such a limiting factor, for example larger UAV’s operated by a military.
These are the kind of numbers an ideal SFM should be able to use under deployment conditions.
Any team building toward SFMs should be able to point to a concrete answer for each of these four core pillars. A generic transformer with a smaller parameter count and a re-sampled input pipeline is not an SFM, whatever the marketing may say in the future.
Use cases
I think SFMs deserve to be named as their own category because the underlying technical problem repeats, almost unchanged, across multiple domains that share nothing in terms of go-to-market or regulatory environment.
Autonomous vehicles and drones. This is something I have already covered, and there is a good reason for it. I think it shows the clearest initial commercial path for SFM’s in the short term. These industries have the best use I can see for multi-sensor fusion (IMU, GPS, vision, LIDAR, barometric) under strict real-time control loops. They also often have severe power/weight budgets and this is close to the purest expression of the SFM problem.
Consumer and smart environments. Sensors that provide heterogeneous data are becoming more prevalent across homes and buildings. Motion sensing, radar, air quality, energy, occupancy, acoustic, and imaging, all reporting asynchronously. Each needs decisions made locally for latency, privacy, and reliability reasons rather than round-tripped to the cloud.
Space and remote or off-grid systems. Satellites and remote infrastructure have a varied amount of power and computational budget, but in some cases face even stricter decision-making constraints than drones or autonomous vehicles.
Healthcare and wearables. Continuous physiological monitoring is asynchronous by nature. Just take heart rate and blood-oxygen levels, these are inherently asynchronous and the cost of a missed or delayed inference could result in a negative patient outcome.
I don't expect one model to run unmodified across a drone, a satellite, and a thermostat. The go-to-market and regulatory environment for each is completely different. What I do expect to carry across all of them is the architecture itself: being multimodal native, doing asynchronous data fusion, reasoning and acting in real-time, and running local first. That's the same pattern that let "predict the next token" become the common capability underneath translation, summarisation, and code generation. Whoever builds the best general solution to that capability has a single foundation model business spanning drones, homes, satellites, and wearables.
Why Now
I think three trends that were each individually true, but not yet convergent, have come together only in the last couple of years, and it's their convergence that makes SFMs feasible now rather than a research curiosity for the 2030s. I say 2030’s here for dramatic effect, not because I think it would have taken us this long to get there. Time has been a fickle thing in the AI space, and it has been very hard to estimate when the next big jump will come. AGI was meant to arrive by the end of 2026, and as I write this, we are only 4 months away from that target date.
The scale of the underlying trend I believe is already visible in the numbers. Estimates put the number of connected IoT devices at roughly 21 billion, with forecasts of more than 40 billion by the mid-2030s as sensors proliferate across buildings, vehicles, and industrial equipment.9,10 Independent market analyses point toward steep growth in edge AI, TinyML, and smart-building software over the next decade. The specific figures vary by methodology and are best read as approximate. Every major estimate points the same way: more sensors, more edge compute, and a widening gap between the data being generated and the AI capable of using it in real-time.
Edge silicon crossed a threshold. Neural accelerators are now common in mid-range microcontrollers and SoCs, at power and cost points that make on-device inference viable for consumer-scale hardware rather than only defence and aerospace budgets. This being said, we are currently in the “RAM-pocalypse” which has made the price skyrocket for many important IC’s and anything that is used in the AI space.
Sensors got radically cheaper and denser. The cost of instrumenting a physical environment with a dozen heterogeneous sensors has fallen enough that "sensor-rich by default" is now a viable product decision for consumer hardware, not just industrial equipment. That means the training data SFMs need is finally being generated at scale, by devices already being deployed for other reasons.
Neuro-symbolic and hybrid architectures are becoming credible. Pure end-to-end deep learning struggles with the safety guarantees and sample efficiency the physical world demands; growing use of approaches that combine learned representations with symbolic constraints and verifiable rules is giving SFMs a more credible path to real-time reasoning that a regulator, an insurer, or a safety engineer could plausibly sign off on.
On their own, none of these three would be enough. Cheap sensors without enough edge silicon just produce data that nothing can act on locally. Edge silicon without better reasoning methods produces fast models that are still black boxes in domains that can't tolerate one. It's the combination of all three, arriving together, that has made these possible.
What SFMs Could Mean for Investors
Before investors, let’s touch on engineers. For them, I think this means the future is closer than it was before. Systems can "think" in real-time and then act, based on real-world data. This is where reality starts to meet sci-fi.
For investors, the question is where value accrues once the category exists. I believe there are four plays worth watching.
The model layer. Whoever trains the best general-purpose SFM model, and makes it easy to adapt across domains, occupies a position similar to the foundation model labs in language and vision. This will be a valuable, but contestable, and margins are likely to compress over time as different players enter the space. However, a first mover advantage might hold out and could present a very valuable opportunity.
The silicon layer. Edge accelerators purpose-built for asynchronous, sparse, low-power inference are a deep hardware moat, but hardware cycles are slow and capital-intensive. I'd expect the winners here to be incumbents who already have a fabrication and distribution advantage.
The data layer. Another area is a company that's already deploying dense, heterogeneous sensor hardware into physical environments at scale, for a business reason independent of SFM (energy savings, safety, automation), and accumulating exactly the messy, asynchronous, real-world data that SFM training requires. Cloud-scale text and image data is now largely spoken for; real-world multimodal sensor data at scale is not, yet.
The full stack layer. For me, this is the most interesting position. A company that not only trains the model, but also can utilise it in the physical world, rather than just licencing it. That way the company can train the model on its own sensor data, whilst generating multiple revenue streams, all whilst having a very strong defensible moat. This is of course the hardest option, but it is the option with the highest upside.
I expect the category to follow the same arc as language foundation models: an initial period where general-purpose backbones get most of the attention, followed by a longer period where the durable value shifts toward whoever controls proprietary deployment data and the distribution to keep collecting it. I'd weight that second phase more heavily than the current headlines tend to.
If SFMs are going to be a category and not just an idea, it needs the building blocks every new foundation model category has required before it. These are: shared benchmarks that actually measure asynchronous, heterogeneous, power-constrained performance rather than repurposed vision or NLP benchmarks; open reference architectures that make the four pillars above concrete rather than aspirational; and a common vocabulary, so that "sensor AI," "edge AI," and "SFM" stop being used interchangeably for very different levels of ambition. That's the gap between "edge AI" as it's mostly practiced today (compressed, single-modality classifiers) and SFMs as I've defined it here. Closing it is a multi-year systems and research problem, not a fine-tuning exercise, and I think it'll be won by teams that treat real-world deployment as the primary source of training signals, not a downstream afterthought to a model trained on public benchmarks.
The physical world produces more data, more continuously, than the text and image datasets that built today's foundation models, and only a very small amount of it is being used to train new models, because the tools built for language and vision models don’t currently support it as well as they could. SFMs are what it could look like to take the foundation model recipe seriously for the data the world actually generates: asynchronous, heterogeneous, real-time, and power-constrained. The technology to build them is ready. The category doesn't have a name yet that's stuck. SFM(s) is my proposal for one.