5.2.1 Voice Control

The voice control interface provides comprehensive voice interaction capabilities, including speech synthesis, speech recognition, audio noise reduction, audio playback, and volume control.

Key Features

Text-to-Speech (TTS)

  • Text-to-speech: Convert text into natural-sounding speech.

  • Multi-language support: Supports Chinese, English, and other languages.

  • Emotional speech: Supports different emotional styles for synthesis.

  • Priority management: Supports multi-level priority control.

Automatic Speech Recognition (ASR)

  • Real-time recognition: Supports real-time speech recognition.

  • Multi-language recognition: Supports Chinese, English, and other languages.

  • Audio stream processing: Supports real-time processing of audio streams.

Audio Processing

  • Real-time noise reduction: Supports real-time audio denoising.

  • Voice activity detection: Supports VAD (Voice Activity Detection).

  • Streaming: Supports streaming of denoised audio.

Audio Playback

  • Audio stream playback: Supports playback of audio data streams.

  • Priority control: Supports playback priority management.

  • Format support: Supports multiple audio formats.

Volume Control

  • Volume adjustment: Supports system volume adjustment.

  • Mute control: Supports mute / unmute.

  • Volume query: Supports querying the current volume.

Volume Control Services

Service Name

Data Type

Description

/aimdk_5Fmsgs/srv/GetVolume

GetVolume

Query volume

/aimdk_5Fmsgs/srv/SetVolume

SetVolume

Set volume

/aimdk_5Fmsgs/srv/GetMute

GetMute

Query mute status

/aimdk_5Fmsgs/srv/SetMute

SetMute

Set mute

  • GetVolume ros2-srv @ hal/audio/srv/GetVolume.srv

    # Get Volume
    # Service: /aimdk_5Fmsgs/srv/GetVolume
    
    # Request
    CommonRequest request            # Request header
    
    ---
    
    # Response
    CommonResponse reponse           # Response header
    uint32 audio_volume              # Current volume (0–100)
    
  • SetVolume ros2-srv @ hal/audio/srv/SetVolume.srv

    # Set Volume
    # Service: /aimdk_5Fmsgs/srv/SetVolume
    
    # Request
    CommonRequest request            # Request header
    uint32 audio_volume              # Target volume (0–100)
    
    ---
    
    # Response
    CommonResponse reponse           # Response header
    uint32 audio_volume              # Current volume (0–100)
    
  • GetMute ros2-srv @ hal/audio/srv/GetMute.srv

    # Get Mute Status
    # Service: /aimdk_5Fmsgs/srv/GetMute
    
    # Request
    CommonRequest request            # Request header
    
    ---
    
    # Response
    CommonResponse reponse           # Response header
    bool is_mute                     # Current mute state
    
  • SetMute ros2-srv @ hal/audio/srv/SetMute.srv

    # Set Mute
    # Service: /aimdk_5Fmsgs/srv/SetMute
    
    # Request
    CommonRequest request            # Request header
    bool is_mute                     # Target mute state
    
    ---
    
    # Response
    CommonResponse reponse           # Response header
    bool is_mute                     # Current mute state
    

Note

SetVolume and SetMute linkage: SetVolume(0) auto-mutes; SetVolume(>0) auto-unmutes. SetMute does not change the volume.

Text-to-Speech Service

Service Name

Data Type

Description

/aimdk_5Fmsgs/srv/PlayTts

PlayTts

Text-to-speech playback

  • PlayTts ros2-srv @ interaction/srv/PlayTts.srv

    # TTS Playback
    # Service: /aimdk_5Fmsgs/srv/PlayTts
    
    # Request
    CommonRequest header
    PlayTtsRequest tts_req  # Embedded request msg
    
    ---
    
    # Response
    CommonResponse header
    PlayTtsResponse tts_resp  # Embedded response msg
    

    Where

    • PlayTtsRequest ros2-msg @ interaction/msg/PlayTtsRequest.msg

      # Embedded request msg
      
      string text                      # Text content
      TtsPriorityLevel priority_level  # Priority level (see TtsPriorityLevel below)
      uint32 priority_weight           # Priority weight (0–99)
      string domain                    # Caller domain
      string trace_id                  # Request trace ID
      bool is_interrupted              # Whether to interrupt broadcasts of the same priority (otherwise queued)
      
      • TtsPriorityLevel ros2-msg @ interaction/msg/TtsPriorityLevel.msg

        # TTS priority level
        uint8 value                      # Priority value
        

        Available TtsPriorityLevel values:

        Level

        Value

        Description

        Usage scenarios

        Emergency safety layer (SAFETY_L10)

        10

        Highest priority

        Safety alerts, emergency notifications

        Warning layer (WARNING_L8)

        8

        High priority

        Hazard alerts and warning messages

        System notice layer (SYSTEM_L7)

        7

        Medium-high priority

        System-level Notice

        Interaction response layer (INTERACTION_L6)

        6

        Medium priority

        User interaction and conversational responses

        Mission execution layer (MISSION_L4)

        4

        Medium-low priority

        Task execution and status broadcasts

        Service layer (SERVICE_L2)

        2

        Low priority

        Proactive services and reminders

        Background service layer (BACKGROUND_L1)

        1

        Lowest priority

        Background services and logging

        Audio playback priority mechanism:

        • This priority system applies to TTS playback (PlayTts).

        • Priority comparison: simple numeric comparison, higher value = higher priority

        • Equal priority behavior: controlled by is_interrupted parameter

          • is_interrupted=true: interrupt current playback

          • is_interrupted=false: queue for sequential playback

        • The playback queue will be cleared when interrupted.

        • The emergency safety level has the highest priority and cannot be interrupted by any other level.

        Attention

        Avoiding voice playback conflicts:

        During runtime, the following modules automatically call PlayTts and may interrupt the user’s voice playback:

        • task_manager (Motion Control Computing Unit, PC1, 10.0.1.40): automatically plays TTS at SAFETY_L10 (e.g. fall / damping-fall / over-temperature) or WARNING_L8 (e.g. navigation interrupted / over-temperature alerts) during fault diagnosis. SAFETY_L10 is the highest priority and cannot be preempted; using SYSTEM_L7 blocks WARNING_L8 and below, but also suppresses fault voice alerts (e.g. fall protection, over-temperature warnings), so the user can no longer perceive those faults via voice.

        • interaction_slave (Interaction Computing Unit, PC3, 10.0.1.42): automatically calls PlayTts at INTERACTION_L6 (priority_level=6, is_interrupted=true) on charging state changes. Use a priority_level higher than 6 (e.g. SYSTEM_L7 or above) to avoid being interrupted.

        To fully stop task_manager’s auto-playback, run the following on the Motion Control Computing Unit (PC1, 10.0.1.40):

        aima em stop-app task_manager
        

        To fully stop interaction_slave’s auto-playback, run the following on the Interaction Computing Unit (PC3, 10.0.1.42):

        aima em stop-app interaction_slave
        

        Note

        After using a high priority or stopping system modules, fault voice alerts (e.g. fall protection, over-temperature warnings) may be suppressed or no longer played. To avoid missing critical fault information, also subscribe to the HDS diagnostic topics (/aima/hds/diag_code_list, /aima/hds/post_proc/signal) to monitor system safety alerts programmatically. See Fault Handling.

        Note

        Audio file playback (PlayAudioFile) uses a different priority system from PlayTts, managed via the audio focus mechanism. See Audio Stream Playback Feature Set for details.

    • PlayTtsResponse ros2-msg @ interaction/msg/PlayTtsResponse.msg

      # Embedded response msg
      string text                      # Response text
      TtsPriorityLevel priority_level  # Priority level
      uint32 priority_weight           # Priority weight
      string domain                    # Caller domain
      string trace_id                  # Request trace ID
      bool is_success                  # Whether the request succeeded
      string error_message             # Error message
      uint32 estimated_duration        # Estimated duration (unit: ms; current firmware always returns 0; call GetTtsDuration to estimate duration)
      uint8 error_code                 # Error code (0 (ERROR_NONE no error), 1 (ERROR_HIGHER_PRIORITY_BUSY preempted by higher priority), 2 (ERROR_NETWORK_DISCONNECTED reserved), 3 (ERROR_UNKNOWN unknown error))
      

Audio File Playback Service

Call the PlayAudioFile service with the audio file path (file_path = parent directory absolute path, file_name = filename) and a priority to trigger playback. The response field reponse.status.value is a status code: 1 = success, 2 = focus request failed, 5 = invalid parameter (unsupported file name/format, inaccessible file, zero-length data, etc.).

Service Name

Data Type

Description

/aimdk_5Fmsgs/srv/PlayAudioFile

PlayAudioFile

Play audio file

  • PlayAudioFile ros2-srv @ hal/audio/srv/PlayAudioFile.srv

    # Play audio file
    # Service: /aimdk_5Fmsgs/srv/PlayAudioFile
    
    # Request
    CommonRequest request            # Request header
    AudioFile file                   # Audio file info (required)
    builtin_interfaces/Time play_stamps  # Optional; specifies the playback time (UTC), defaults to immediate playback
    
    ---
    
    # Response
    CommonResponse reponse           # Response header
    

    Note

    reponse.status.value is always SUCCESS(1) regardless of whether focus is granted. Check focus_response.focus_gain (true = granted, false = not granted).

    • AudioFile ros2-msg @ hal/audio/msg/AudioFile.msg

      string pkg_name        # Required; identifies the caller source
      string file_name       # Required; file name
      string file_path       # Required; file path (uses the system default path when not set; must not end with the file name)
      AudioInfo info         # Required for PCM files; not used for WAV files (sample rate and channel count are parsed from the WAV header)
      uint32 priority        # Required; priority (1–10, default 6)
      uint32 priority_weight # Optional; (1–100) final priority = priority + priority_weight%
      

    Notes:

    • Audio files must be raw PCM (.pcm) or WAV (.wav; encoding must be PCM or WAVE_FORMAT_EXTENSIBLE with a PCM SubFormat). MP3 and other formats are not supported.

    • Audio data must be 16-bit signed integer PCM (S16LE), with non-zero data length. Sample rate and channel count are unrestricted; the system automatically resamples to 16 kHz / mono. For best results, use 16 kHz / mono.

    • For PCM files, set sample_rate and channels via the info field; WAV files do not use the info field — the sample rate and channel count are parsed automatically from the WAV header.

    • file_path is the parent directory’s absolute path (e.g. /var/tmp/my_audio), and file_name is the file name (e.g. test.wav).

    • Audio files must be stored on the interaction compute unit (PC3, 10.0.1.42), not the development compute unit (PC2).

    • The audio folder and all its parent directories must be readable by all users (a subdirectory under /var/tmp/ is recommended).

Audio Stream Playback

Provides raw audio stream playback support

Warning

Focus management notes

When pushing audio data via the /aima/hal/audio/playback topic, the hal_audio subscriber does not automatically request focus. Users must call the RequestAudioFocus service themselves before playback; otherwise, it may conflict with other playback sources.

Only the PlayAudioFile service manages focus automatically.

Service Name

Data Type

Description

/aimdk_5Fmsgs/srv/RequestAudioFocus

RequestAudioFocus

Request audio playback focus

/aimdk_5Fmsgs/srv/AbandonAudioFocus

AbandonAudioFocus

Release audio playback focus

/aimdk_5Fmsgs/srv/StopAudioPlay

StopAudioPlay

Stop audio playback

Topic Name

Data Type

Description

QoS

Frequency

/aima/hal/audio/playback

AudioPlayback

Audio stream playback

RELIABLE+VOLATILE

User (on-demand)

/aima/hal/audio/focus_response

FocusResponse

Audio focus change events

RELIABLE+TRANSIENT_LOCAL

Event-driven

/aima/hal/audio/play_state

PlayStateChange

Audio playback state events

BEST_EFFORT+TRANSIENT_LOCAL

Event-driven

  • RequestAudioFocus ros2-srv @ hal/audio/srv/RequestAudioFocus.srv

    # Request audio playback focus
    # Service: /aimdk_5Fmsgs/srv/RequestAudioFocus
    
    # Request
    CommonRequest request # Request header
    
    FocusRequester focus_requester # Focus request info
    
    ---
    
    # Response
    CommonResponse reponse # Response header
    
    FocusResponse focus_response # Request result
    

    Note

    reponse.status.value is always SUCCESS(1), indicating only that the request was processed. Check focus_response.focus_gain (true = granted, false = not granted).

    • FocusRequester ros2-msg @ hal/audio/msg/FocusRequester.msg

      string pkg_name # Playback source identifier
      
      uint32 priority # Priority (1–10, default 6)
      
      uint32 priority_weight # Weight (optional); breaks ties within same priority level
      
    • FocusResponse ros2-msg @ hal/audio/msg/FocusResponse.msg

      string pkg_name # Playback source identifier
      
      bool focus_gain # Focus grant result
      
  • AbandonAudioFocus ros2-srv @ hal/audio/srv/AbandonAudioFocus.srv

    # Release audio playback focus
    # Service: /aimdk_5Fmsgs/srv/AbandonAudioFocus
    
    # Request
    CommonRequest request # Request header
    
    FocusRequester focus_requester # Focus request info
    
    ---
    
    # Response
    CommonResponse reponse # Response header
    
    FocusResponse focus_response # Request result
    

    Note

    reponse.status.value is always SUCCESS(1), indicating only that the request was processed. Check focus_response.focus_gain to determine whether focus was released.

    • FocusRequester and FocusResponse are defined as above >>

    Warning

    AbandonAudioFocus Notes

    When calling AbandonAudioFocus, the pkg_name, priority, and priority_weight fields in focus_requester must match exactly the values passed in the corresponding RequestAudioFocus call; otherwise the focus cannot be matched and released.

    Focus preemption rules: A challenger with higher or equal priority preempts the current focus holder, and a focus-loss notification is published via the /aima/hal/audio/focus_response topic. When the same pkg_name repeatedly requests focus, the request is silently updated without publishing a focus-loss notification.

  • StopAudioPlay ros2-srv @ hal/audio/srv/StopAudioPlay.srv

    # Stop audio playback
    # Service: /aimdk_5Fmsgs/srv/StopAudioPlay
    
    # Request
    CommonRequest request # Request header
    
    ---
    
    # Response
    CommonResponse reponse # Response header
    

    Calling this service immediately clears the playback queue and audio buffer, stopping any audio currently playing.

    Warning

    This service stops all audio currently playing; consider the risk of interrupting safety prompts such as alert announcements.

  • AudioPlayback ros2-msg @ hal/audio/msg/AudioPlayback.msg

    # Audio stream playback
    # Topic: /aima/hal/audio/playback
    
    builtin_interfaces/Time stamps # Timestamp
    
    AudioInfo info # Audio format
    
    AudioData data # Audio data
    
    string pkg_name # Playback source identifier
    
    string token_id # (Optional) changing token_id clears the playback buffer (used to interrupt current playback)
    

    Note: When a message with a different pkg_name arrives, hal_audio flushes all old data (clearing the queue and ring buffer), publishes PLAYER_STATE_STOPED, and then starts playing the new data. Messages with different pkg_name values do not coexist.

    • AudioInfo ros2-msg @ hal/audio/msg/AudioInfo.msg

      uint8 channels # Number of channels (required when playing PCM files and audio streams)
      
      uint32 sample_rate # Sample rate [Hz] (required when playing PCM files and audio streams)
      
      uint32 size # Reserved
      
      string sample_format # Reserved
      
      string coding_format # Reserved
      

    Note: size, sample_format, and coding_format are reserved fields; the system always processes data as S16LE PCM. For WAV file playback, the info field is not read at all (sample rate and channel count are parsed automatically from the WAV header).

    • AudioData ros2-msg @ hal/audio/msg/AudioData.msg

      uint8[] data
      
  • FocusResponse ros2-msg @ hal/audio/msg/FocusResponse.msg

    # Audio focus change events
    # Topic: /aima/hal/audio/focus_response
    
    string pkg_name # Playback source identifier
    
    bool focus_gain # Focus grant result
    
  • PlayStateChange ros2-msg @ hal/audio/msg/PlayStateChange.msg

    # Audio playback state events
    # Topic: /aima/hal/audio/play_state
    
    string pkg_name # Playback source identifier
    
    PlayStateType state # Playback state (state.value: 0 (PLAYER_STATE_CLOSED closed), 1 (PLAYER_STATE_PLAYING playing), 2 (PLAYER_STATE_STOPED stopped))
    

Note

PlayStateType state transitions:

  • CLOSED(0): Published only once at module initialization, never appears again afterward.

  • PLAYING(1) ↔ STOPPED(2): PLAYING is pushed when playback starts; STOPPED is pushed when playback ends, is interrupted, or the buffer is exhausted. The two states alternate.

MIC Audio Stream Capture Topic

Supports receiving real-time VAD (Voice Activity Detection) events on denoised audio and the corresponding audio stream, as well as raw audio stream capture.

Topic Name

Data Type

Description

QoS

Frequency

/agent/process_audio_output

ProcessedAudioOutput

VAD audio capture

BEST_EFFORT+TRANSIENT_LOCAL

Event-triggered, cached data for voice recognition would be sent in a burst at start of VAD event, then would update at ~25Hz

/aima/hal/audio/capture

AudioCapture

Raw audio capture

RELIABLE+VOLATILE

33.3Hz

  • ProcessedAudioOutput ros2-msg @ interaction/msg/ProcessedAudioOutput.msg

    MessageHeader header  # Message header (not enabled)
    
    uint32 stream_id  # Audio stream ID (0: built-in mic, 1: external mic; matches SetMicSource service ID. Audio data is published via the corresponding stream_id)
    AudioVadStateType audio_vad_state  # VAD state (audio_vad_state.value: 1 (AUDIO_VAD_STATE_BEGIN speech start), 2 (AUDIO_VAD_STATE_PROCESSING speech processing), 3 (AUDIO_VAD_STATE_END speech end))
    uint8[] audio_data # Audio data (PCM, 16kHz/16bit/1ch)
    MessageHeader header  # Message header
    
    uint32 stream_id  # Audio stream ID (0: onboard mic, 1: external mic; matches the SetMicSource service ID. Whichever mic is active publishes with its corresponding stream_id)
    AudioVadStateType audio_vad_state  # VAD state (audio_vad_state.value: 0 (AUDIO_VAD_STATE_NONE none), 1 (AUDIO_VAD_STATE_BEGIN speech start), 2 (AUDIO_VAD_STATE_PROCESSING speech processing), 3 (AUDIO_VAD_STATE_END speech end))
    uint8[] audio_data  # Audio data (PCM, 16 kHz / 16 bit / 1 ch)
    

Audio stream format:

  • Sample rate: 16 kHz

  • Bit depth: 16-bit (S16LE)

  • Channels: mono

  • Encoding: PCM

The above is the system’s fixed output format. When playing an audio stream via /aima/hal/audio/playback, the system automatically resamples the input data to this format, so the input sample rate and channel count are unrestricted, but the bit depth must be 16-bit.

Attention

The wake word required to activate VAD (since v0.9):

  • In default mode (built-in interaction ON), always say the wake word before target voice, as VAD only keeps activated for a short while.

  • In only_voice mode (built-in interaction disabled), VAD keeps activated for long once woken by the wake word. No more wake words needed later, all voice detected later on would be captured as VAD streams.

  • AudioCapture ros2-msg @ hal/audio/msg/AudioCapture.msg

    # Raw audio capture
    # Topic: /aima/hal/audio/capture
    
    builtin_interfaces/Time stamps
    
    uint8 mic_channels # Number of microphone channels
    
    uint8 ref_channels # Number of reference (echo-cancellation) channels
    
    AudioInfo info # Audio format (info.channels = mic_channels + ref_channels, i.e. total channels)
    
    AudioData data # Audio data
    
    string pkg_name # Audio source (always empty string in current firmware)
    

    AudioInfo definition >> AudioData definition >>

Note

AudioCapture data layout: Multi-channel audio in data is interleaved sample-by-sample (int16_t) — all channel samples at the same instant are laid out sequentially. For example, 2 channels: (ch0_sample, ch1_sample), (ch0_sample, ch1_sample), ... (each pair of parentheses contains all channel samples at the same instant). Channel composition:

  • Built-in mic: 6 channels (4 mic + 2 reference)

  • External mic: 2 channels (1 mic + 1 reference)

Microphone Control Service

Service Name

Data Type

Description

/aimdk_5Fmsgs/srv/GetMicSourceRequest

GetMicSourceRequest

Query the current MIC device

/aimdk_5Fmsgs/srv/SetMicSourceRequest

SetMicSourceRequest

Switch the MIC device

  • GetMicSourceRequest ros2-srv @ interaction/srv/GetMicSourceRequest.srv

    # Query current MIC device
    # Service: /aimdk_5Fmsgs/srv/GetMicSourceRequest
    
    # Request
    CommonRequest header
    
    ---
    
    # Response
    CommonResponse header
    uint32 mic_source  # MIC source (0: built-in mic, non-zero: external mic)
    
  • SetMicSourceRequest ros2-srv @ interaction/srv/SetMicSourceRequest.srv

    # Switch MIC device
    # Service: /aimdk_5Fmsgs/srv/SetMicSourceRequest
    
    # Request
    CommonRequest header
    uint32 mic_source  # 0: built-in mic, 1: external mic
    
    ---
    
    # Response
    CommonResponse header
    

Programming Examples

For detailed programming examples and code descriptions, see:

Safety Notes

Warning

Voice playback limitations

  • The TTS service uses a priority system; avoid starting multiple speech playbacks at the same time.

  • Higher-priority speech will interrupt lower-priority speech; configure priorities carefully.

  • Check the current playback state before starting new speech.

Caution

As standard ROS DO NOT handle cross-host service (request-response) well, please refer to SDK examples to use open interfaces in a robust way (with protection mechanisms e.g. exception safety and retransmission)

While the robot is in Stable Standing Mode or Locomotion Mode, DO NOT launch ROS nodes in rapid bulk (no more than 2 nodes per second is recommended), as a large number of nodes joining DDS discovery within a short period causes communication congestion and degrades motion control real-time performance, which may cause the robot to lose balance and fall

Note

Best Practices

  • Choose appropriate priority levels to avoid interfering with important announcements.

  • Implement monitoring and exception handling for speech playback.

  • Implement a playback queue for speech management.

  • Ensure the audio format and sample rate meet the requirements.

  • The receive queue (QoS depth) of VAD should be large enough

  • Never forget wake words when using VAD