
The core idea of MARII is to pick two of the most important mediums of human communication, voice and gestures, and combine them to create a significantly better interaction experience than the current state of text-based prompting. By delivering context over two different modalities simultaneously, the user can express intent much faster and more naturally than typing.
For example, instead of typing "Move ship 1 to the windfarm located at coordinates 54.32°N, 10.14°E", the user can simply point at the ship, point at the windfarm, and say "Move this ship here." The gesture provides the spatial reference, and the voice provides the action and intent. This division of labor is precisely how humans naturally communicate about physical spaces.
.
.
.

MARII is a platform that studies how humans and AI agents in maritime environments work together under pressure.
Operators command a fleet of vessels using hand gestures and voice, on a screen.
Behind the scenes, an AI orchestrator interprets human intent in real time and decides how much control to hand back.
By measuring cognitive load and response times across rescue, resource-allocation, and crisis scenarios, MARII asks one question:
When the stakes are high at sea who should be in charge? The human, the machine, or the space between them?
.
.
.
Video Capture & 3D Landmark Tracking
.
.
30 FPS @ 480×300 RGB
21 3D Landmarks (x,y,z)
Hand Skeleton Locked
MARII captures raw video from a standard webcam, downscaling frames to 480 × 300 pixels to optimize memory bandwidth. Using MediaPipe's neural landmarker, the system extracts 21
three-dimensional joint coordinates (x,y,z) at 30 FPS. Vector dot-product angle calculations evaluate joint flexion in real time, locking onto the operator's hand skeleton with zero frame drops.
.
Gesture Vocabulary & Actions
.
To eliminate cognitive fatigue, MARII relies on a compact, natural gesture vocabulary: Point (spatial selection), Grab (entity lock), Thumbs Up / Down (AI route confirmation), Wave (canvas reset), and Draw (route tracing). Each pose generates a 5-bit binary finger vector, pairing spatial intent directly with spoken voice commands.
.
Camera to Map Projection
.
Operating on planar projective geometry, MARII applies a calibrated 3 × 3 Planar Homography Matrix (H) to translate raw camera pixel coordinates directly onto geographic latitude and longitude (lat,lon) on the Leaflet.js map display. With an execution latency under 0.4 ms, operators experience fluid, lag-free cursor
pointing across maritime operational charts.
.
.
.
.
.
MARII (Maritime Intelligent Interface) transforms human-AI communication by replacing text prompting with real-time 3D hand gesture tracking and natural voice commands. By eliminating coordinate typing and back-and-forth prompt iterations (ping-pong prompting), MARII empowers operators to command complex maritime fleets through intuitive pointing and speaking—delivering sub-second, zero-typing execution.
.
.
.
Tomas Herrero Valero
Universitat Politècnica de València
tomas.herrero.valero@tha.de
Felix Baeuml
Technische Hochschule Augsburg
felix.baeuml@tha.de
Piotr Grzelak
Politechnika Lodzka
piotr.grzelak@tha.de
.


