Picture album

Front view of the Sansheng Digital Human All-in-One Machine
01-Front
Image showing the back interface of the Sansheng Digital Human All-in-One Machine
01-Back

I. About the Product

1.1 Product Introduction

This product is a voice-interactive digital human system suitable for private deployment in enterprises. It is applicable to intelligent voice interaction and data query/question-answering in enterprises, schools, and government units. Developed based on a large-scale inference model, speech recognition model, speech synthesis model, and workflow engine, it can flexibly switch between public network and local deployment models to meet the personalized data processing needs of enterprises. It focuses on business value, is lightweight and customizable, and can be quickly deployed.

1.2 Character Portrayal

It supports customizing your own character image and offers more than ten images to choose from in the image library.

1.3 Core Advantages

The core difference between us and other digital human manufacturers:

  • Multi-terminal use: Not limited to all-in-one machines, it can be used flexibly on multiple terminals such as Windows computers, Android tablets (class signs), web pages, official accounts, and mini programs.
  • Private Question Answering: It can flexibly connect with internal enterprise data through interfaces, database queries, etc., allowing digital humans to answer exclusive content and enabling flexible querying of reports, videos, and other content.
  • Localized deployment: Supports localized private deployment within enterprises, running on the intranet (all large models and data are local).

1.4 Software Function List

Serial NumbercategorySub-itemillustrate
1Character customization1.1 Character CloningGreen screen shooting: Digital human cloning based on green screen shooting videos provided by the target person; AI synthesis: Generating a person's image based on a person's photo; Supports various actions such as waiting, standing still with hands down, speaking with hands together, speaking with one hand, and speaking with both hands.
1.2 Character Voice CloningIt supports 3-second rapid replication, providing 3-10 seconds of real human voice recordings to quickly replicate the voice characteristics of the person.
1.3 Background SettingsProvides options for setting and changing the background image of the digital human.
2Voice technology2.1 Voice wake-upIt supports naming digital humans, and the digital human will actively wake up and engage in conversation when it hears someone calling its name.
2.2 Speech RecognitionBased on a large speech model, this real-time, end-to-end speech recognition framework boasts industry-leading word error rate and concurrent inference speed. It supports dialect recognition and can identify more than 10 national languages.
2.3 Speech SynthesisSpeech synthesis based on a large speech model, providing multiple voice options. Audio algorithms: AEC, ANC, AGC.

1.5 Digital Human All-in-One Machine

A smart interactive terminal device that integrates a camera and microphone and is ready to use right out of the box.

II. Application Scenarios

2.1 Touchscreen All-in-One PC

This vertical touchscreen all-in-one machine is suitable for scenarios such as lobbies and exhibition halls, allowing users to interact directly by touch or voice.

2.2 Electronic Picture Frame

The digital human interactive screen, shaped like an art frame, can be integrated into home and office environments, allowing users to converse with the digital human within the electronic frame through voice-activated question and answer.

2.3 Exhibition Hall Guide

A digital human-based explanation solution suitable for scenarios such as school history museums, corporate showrooms, and museums, supporting large-screen integrated display.

2.4 Industrial Data Query

This digital human, designed for industrial scenarios, allows users to query production data, equipment status, and other information via voice interaction.

2.5 Windows PC

projectSpecification
Installation methodProvides exe installer
Applicable terminalsPodium computer/Touchscreen all-in-one machine
CPUi3 and above processors
Typical scenariosClassroom podium computer, digital human interdisciplinary knowledge quiz

2.6 Electronic Class Sign

projectSpecification
Display screenIPS LCD display
size18.5 inches
resolution1920 × 1080
chipRK3288 Quad-core 1.8GHz
Memory2 GB
flash memory16 GB

2.7 Industrial Robot Dog

A quadrupedal bionic robot dog, a mobile digital human interaction platform, capable of moving through complex terrain and interacting with users via voice.

2.8 Large-scale conferences

The extra-large interactive screen solution is suitable for digital human display and interaction in scenarios such as large conferences and lecture halls.

III. Character Cloning

3.1 Green Screen Shooting

Digital human cloning is performed based on green screen videos provided by the target individual. Recording must be done in a professional green screen environment. After submitting high-definition video footage, image training can typically be completed within 1-2 business days.

3.2 Photo Combination

AI-generated avatars based on photographs can quickly create digital human appearances without the need for specialized equipment.

3.3 Sound Synthesis

Supports rapid voice feature replication in 3 seconds. Only 3-10 seconds of real-person speech recording is required to replicate the person's voice characteristics for use in digital human speech synthesis.

IV. Front-end Digital Human

4.1 Digital Human Q&A

4.1.1 Text Response

The digital human replies to user questions in plain text format, which is suitable for general knowledge-based question-and-answer scenarios.

4.1.2 Answer with pictures and text

When a question explicitly requires a "text and images" response, the digital human can display both text and images in the reply.

4.1.3 Video Answers

When the question explicitly requires a "video answer", the digital human can access and play the video materials already configured in the backend (the backend resources must have the corresponding video materials).

4.2 Interface Switching

4.2.1 Multi-mode switching

You can freely switch between voice input and text input modes.

4.2.2 Multiple Background Switching

Supports switching the background image of the digital human.

4.2.3 Multi-resolution switching

It offers multiple resolutions, including 1K, 2K, and 4K, suitable for both regular devices and large-screen all-in-one machines.

4.2.4 Switching between multiple digital human figures

Windows: Select the digital human you want to run through the digital human application; Android: Click the dot in the upper left corner of the interface to bring up the digital human switching interface.

4.2.5 Switching between landscape and portrait modes

Landscape mode is suitable for digital human lectures and company knowledge presentations; portrait mode is suitable for large-screen all-in-one machines. Click the menu on the right to switch.

4.3 Switching between dialogue modes

4.3.1 Wake-up Mode

The system defaults to wake-up mode, which automatically wakes up and begins a conversation when the digital human hears someone call its "name".

4.3.2 Text Input Mode

Click the "Input Mode" menu, and a text input box will pop up. Type in the relevant content to talk to the digital human.

4.3.3 Mobile Phone Dialogue Mode

Click the "Voice Input" menu, then press and hold the button below to converse with the digital human via voice.

4.4 APP Management

4.4.1 App Upgrade

Supports online upgrades of the digital human client via the APP.

4.4.2 App Authorization

It supports app authorization management, allowing control over the usage permissions of different users.

V. Backend Management System

5.1 System Operation

5.1.1 Client Download

Download the client installation packages for each platform from the management backend.

5.1.2 Multi-digit Human Management

We provide one to dozens of digital avatars for different clients, which are managed uniformly through the avatar information interface.

5.1.3 Digital Human Renaming

The digital human's name is the wake word; when the digital human hears this name, it begins to converse.

Special Note: This technology uses an offline voice wake-up SDK, boasting a low false positive rate and industry-leading recognition accuracy. However, it is recommended that the digital human's name have at least four characters (e.g., "Teacher Xiao Wang"). If only two characters are used (e.g., "Xiao Wang"), the digital human may still be falsely awakened when it hears "Xiao Wang". After renaming, the project server must be restarted (please contact the administrator to restart).

5.1.4 Dialogue Log Inquiry

You can filter by different digital users to view recent conversation records.

5.1.5 User System Authorization

Different menu permissions and data permissions are granted to different types of users.

5.2 Document Knowledge Base Management

📌 Regular users can simply upload files; no adjustments to the segmentation strategy are required.

5.2.1 File Management

Select local file upload (drag and drop supported), the system will split the file according to the default strategy: chunk_size=500, chunk_overlap=50, separator=['\n\n'].

5.2.2 Segmented Management

It supports viewing and manually adjusting segments of uploaded documents.

5.3 QA Knowledge Base Management

5.3.1 Question-and-answer pair management

Manually maintain standard question-and-answer pairs, suitable for high-frequency, precise response scenarios.

5.3.2 Precise Q&A Responses

It supports configuring precise matching rules to ensure that a preset standard answer is returned for a specific question.

5.4 Large Model Related

5.4.1 Model Selection

It supports flexible switching between the public network large model and the local deployment model.

5.4.2 Model Training

It supports fine-tuning and training models based on enterprise private data.

5.5 Digital Human Workflow

5.5.1 Workflow Management

Visually orchestrate the digital human's question-and-answer process, supporting drag-and-drop node configuration.

Node NameFunction Description
5.5.2 Starting NodeWorkflow start node (automatically created, cannot be copied or deleted). Configuration: Opening message is empty; automatically retrieves the current time (yyyy-mm-dd hh:mm); stores the most recent n chat records, the number of records can be customized.
5.5.3 Input NodeIt receives the content entered by the user in the dialog box and stores it in the user_input variable, requiring no additional configuration.
5.5.4 Output NodeDefine the final response content returned to the user.
5.5.5 Large Model NodesCall LLM to perform inference and generate responses.
5.5.6 Assistant NodeConfigure system-level prompt commands and assistant behaviors.
5.5.7 QA Knowledge Base Retrieval NodeRetrieve matching question-answer pairs from the QA knowledge base.
5.5.8 Document Knowledge Base Question and Answer NodeRetrieve relevant content from the document knowledge base to answer the question.
5.5.9 Conditional Branch NodeDetermine the flow of the process based on conditions to implement complex dialogue logic.

5.6 System Tools

Provides auxiliary tools and functions related to system operation and maintenance.

VI. SaaS side: Creating a new digital human

stepoperateillustrate
first step6.1 Registered UsersRegister an account on the SaaS platform.
Step 26.2.1 Creating a new workflowCreate a dialogue flow for digital humans.
6.2.2 Creating a new knowledge baseUpload documents or configure QA question-and-answer pairs.
Step 36.3.1 Creating a new character imageChoose or customize the appearance of your digital human.
6.3.2 Association ConfigurationEnsure that the workflow and knowledge base of this person are properly linked.
Step 46.4.1 Obtaining the Serial NumberClick the small circle in the upper left corner of the device to check the device serial number.
6.4.2 Device BindingEnter the serial number in the "Bind Device" section of the management system. After restarting the app, the device will display the newly created digital human, whose answers are based on a dedicated knowledge base.

VII. Privatization Deployment

7.1 System Server Configuration

matterillustrate
7.1.1 Digital Human All-in-One MachineHardware devices that integrate a touchscreen, camera, and microphone, ready to use right out of the box.
7.1.2 Large Model Computing Server (Optional)For running large models locally, a GPU server (NVIDIA RTX 3060 or higher) with 16GB+ of RAM and 50GB+ of available hard disk space is recommended, depending on the concurrency and model size.

7.2 Project Data Compilation

Data typeillustrate
7.2.1 Character appearance and voiceProvide a green screen video or high-definition photo of the target person, along with a 3-10 second audio sample.
7.2.2 QA Question and Answer CompilationCompile common enterprise questions and standard answers, and import them into the QA knowledge base.
7.2.3 Data Processing (RAG)Organize enterprise documents (supports PDF, Word, TXT, Markdown, and other formats), and the system automatically parses and generates enhanced RAG search results.

7.3 Data Integration

  • 7.3.1 Integration with QA Knowledge Base: Import the compiled QA question and answer pairs into the system.
  • 7.3.2 Integration with document knowledge base: Upload enterprise documents, and the system will automatically segment and index them.
  • 7.3.3 Database Query Integration: By integrating with the enterprise's internal database, the digital human can query business data (such as reports, work orders, etc.) in real time.

VIII. Precautions

8.1 Timed power on/off

Set timers for power on/off via tablet devices, such as turning off at 10 PM and turning on at 8 AM, to ensure stable operation and save energy.

微信二维码 WeChat
WhatsApp二维码 WhatAPP
Online consultation