Robots don't need better motors. They need intelligence.
Humanoid and quadruped hardware has quietly matured. Cheap, capable, developer-friendly platforms already exist. What is still missing is the layer that turns motion into awareness, and awareness into useful behaviour. That is where AlifZetta comes in.

Today's robots are choreographed, not aware.
A quadruped that can dance, back-flip, and open a door still cannot recognise your grandmother's voice, notice she has not moved in an hour, or offer to call the neighbour. The mechanical problem is largely solved. The intelligence problem is wide open.
The industry has spent a decade optimising motors, joints, batteries, gait control. The result is genuinely impressive hardware, at a price that would have been science fiction in 2018. But hand any of these platforms to a hospital ward, a family home, or a warehouse floor, and they need a human choreographer to be useful. Perception is bolted on. Conversation is a script. Autonomy is a demo.
AlifZetta Robotics starts from the other end. The hardware is fine. The intelligence layer that turns it into a useful teammate — that is what we are building, on the same sovereign, CPU-native, cited-every-answer stack that runs behind axz.si.
Four capabilities. One intelligence.
We are not selling four different products. We are teaching one robot four things at once, and then getting out of the way.

Perception
Recognise faces, objects, gestures, and postures on-device. The vision model does not phone home. Faces stay in the family home; nothing about the room ever leaves the machine unless the operator explicitly asks.

Conversation
Voice in, voice out, in the language of the room — Nepali, Hindi, English, or the local dialect the operator trained on. Every reply is anchored in the AlifZetta substrate, so the robot will not confidently make up a medication schedule or a fact about your grandfather.

Gesture
Mirror when useful (a namaskar, a wave, a nod). Respond when addressed (turn to face, tilt to listen). Deliberately restrained — a robot that gestures well is welcome; a robot that gestures constantly is uncanny.

Autonomy
Short-horizon behaviours the operator actually asks for — fetch the phone from the table, check on grandmother in the sitting room, patrol the perimeter until 6 a.m., wake me if anything moves. No open-ended promises. Bounded scope, honest failure modes.
Every one of these capabilities runs on the same CPU-native inference stack described in the AlifZetta whitepaper. No external LLM at inference time. No cloud round-trip for the robot to recognise a face or answer a question.
Unitree G1 humanoid.
We picked a hardware partner that had already solved the hard mechanical problems, at a price that lets us iterate every week instead of every quarter.
Unitree Robotics G1

A general-purpose humanoid with a mature SDK, a large developer community, and hardware capable enough to serve as the physical substrate for real perception-and-behaviour experiments. Unitree makes the body; AlifZetta makes the brain.
Reinventing the mechanics would delay the intelligence layer by two years. The G1 lets us focus on the actual open problem: how a robot understands the room it is in, the person it is with, and the task it has been given — and then behaves like a teammate rather than a demo.
We credit the hardware. We are proud of the intelligence.
Why we ship what is working, not what is polished.
Innovation, to me, is about connecting technologies that were not necessarily designed to work together — and seeing what emerges. R&D is not about producing a perfect demo. It is about discovering possibilities, accepting the failures, and learning what the next iteration should look for.
The AlifZetta robotics work is happening in the open. Video from the lab, failure modes we hit and fixed, the specific behaviours we can and cannot reproduce this week — all of it is on the founder's LinkedIn feed and, over time, in the blog. When it works, we say so. When it does not, we say that too, and credit the platform that let us try.
Recent moments from the bench.
Three snapshots from the R&D bench — good weeks, hard weeks, and the moment a bounded behaviour finally worked. Click through for the full LinkedIn write-up.



Where a Nepal-anchored robot earns its keep.
We do not need to solve every application at once. We need to solve two or three well, in operators' actual environments, in Nepal first.
Family homes
Mobility check-ins, medication reminders, fall detection, companionship. The robot notices the grandmother has not moved since breakfast and calls the neighbour before it becomes an emergency.
Ward assistance
Delivery of supplies between the nurse station and patient rooms, patient greeting and check-in, walking companion for post-surgery mobility drills. Frees ward staff for the work that requires a human hand.
First look
Post-earthquake rubble entry where sending a human is unsafe. Post-GLOF flood-zone survey. The robot goes in, the operator watches and listens, decisions are made with sight-and-sound the responder would not otherwise have.
Warehouse & security patrol
Bounded-loop patrol at night, anomaly detection, escalation to a human on call. The robot does not replace the guard; it lets one guard cover the ground that used to take three.
Zero external LLM. Zero cloud round-trip for perception.
Everything the robot does — recognising a face, understanding a sentence, deciding to turn — runs on the machine in front of you. No OpenAI, Anthropic, or Google API is called to answer who is at the door. No cloud service is queried to decide whether the person on the floor has moved in the last five minutes. That is not a marketing line; it is an operating constraint we designed the intelligence layer around.
Two consequences follow.
- Privacy is structural. Faces, voices, and behaviour in the room stay in the room. The operator opts in explicitly for any telemetry that leaves the device — nothing leaks by default.
- The robot works offline. Rural deployments, post-disaster zones, and clinics with unreliable connectivity are the target environment, not the exception. The intelligence does not degrade when the network does.
The technical foundation is the same three-layer stack that powers NEXUS, PRISM, and every answer on demo.axz.si — adapted for embodied inference on modest hardware.
Bring a workload. We bring the intelligence.
We are looking for operators and institutions that have a real task a robot could take on — a hospital ward, a family carer network, a warehouse manager, a municipal disaster office. Bring the problem. We stand up an intelligence layer against it.
R&D collaboration inquiry.
Email padam@axz.si with the task, the environment, and the hardware you already have (or want to try). First reply within 24 hours.
Email the Founder Follow on LinkedIn