Price: 0
Number of applications: 4
13.11.25 (inclusive)
contractual
Idea
ICT tasks
Robotics
Intelligent control systems
Purpose and description of task (project)
Goal: To develop a Windows software utility that captures audio streams (incoming and outgoing audio) from the sound card level of a personal computer for further use in streaming speech recognition (STT) systems. Task description: Create a solution that: 1. Intercepts the audio signal from the local microphone and the sound coming into the headphones/speakers, without interfering with network settings; 2. Generates two media streams (incoming and outgoing) with minimal delay; 3. Provides access to streams via API or local interface for integration with online speech recognition services; 4. Works with popular VoIP applications without requiring their modification; 5. Ensures stable operation and low resource consumption; 6. Does not use port mirroring technology or interference with network equipment.
Software/ IS
In a number of call centers, operators use exclusively softphones that work through third-party VoIP applications to receive and process calls. This approach makes it difficult to access media streams (incoming and outgoing audio) when performing online speech recognition (STT) or real-time conversation analysis. Traditional methods of obtaining audio through network mirroring (port mirroring) or specialized equipment are unacceptable due to the complexity of implementation and limitations of the infrastructure.
The result of the development will be a software utility for Windows OS that provides: 1. Capture audio streams (incoming and outgoing audio) directly from the PC's sound card, regardless of the VoIP application used (for example, Zoiper, MicroSIP, Linphone, Teams, Zoom, etc.). 2. Formation of two separate streams - incoming and outgoing - for subsequent transmission to speech recognition services (STT) with streaming processing; 3. Real-time operation, without significant delays and degradation of sound quality; 4. Compatibility with external STT APIs (Google, Whisper, Yandex SpeechKit, Azure Speech, etc.); 5. No need to configure network equipment or configure VoIP applications; 6. Simple implementation - installation and configuration at the operating system level.
Danchenko Maxim