Generating Instrument Sounds Aligned with Video via Human Body Keypoints : A Deep Learning Approach to Multimodal Audio-Visual Synthesis

Language: English

Published by GRIN Verlag, 2026

3389179836 / 9783389179833

  • Softcover
  • New
See all details

Seller: AHA-BUCH GmbH, Einbeck, GermanyAHA-BUCH GmbH

5-star seller

AbeBooks seller since August 14, 2006

View this seller's items
Softcover

Condition: New

£ 16.74

£ 51.77 shipping 
Ships from Germany to U.S.A.

Quantity: 1 available

Add to basket
Free 30-day returns

Item description from seller

Druck auf Anfrage Neuware - Printed after ordering - Research Paper (undergraduate) from the year 2026 in the subject Computer Science, grade: Good, , language: English, abstract: Historical video archives and recordings from the past often suffer from degraded or completely missing audio tracks due to deterioration of storage media, recording limitations of the era, or loss during archival processes. Similarly, silent films and performance documentation may lack synchronized sound entirely. Emerging generative artificial intelligence techniques have demonstrated the potential to reconstruct missing audio content by analyzing visual information alone-a capability particularly valuable for restoring cultural heritage materials and historical performance recordings. However, when applied to complex activities such as musical instrument performance, existing methods have shown limited accuracy in capturing the nuances of sound production. Prior research has established that SpecVQGAN architectures combined with Transformer-based mechanisms can improve video-to-audio generation. This work introduces an enhanced model that augments SpecVQGAN by incorporating human skeletal pose features, specifically designed to elevate the quality of generated musical instrument sounds. Through comprehensive evaluation using both subjective user studies and objective quantitative metrics, we demonstrate that the proposed framework significantly outperforms existing approaches in reconstructing authentic instrumental audio from archival and silent performance videos.

Seller Inventory # 9783389179833

Title
Generating Instrument Sounds Aligned with Video via Human Body Keypoints : A Deep Learning Approach to Multimodal Audio-Visual Synthesis
Author
Yasuyuki Tahara
Publisher
GRIN Verlag
Publication year
2026
Condition
Neu
Binding
Taschenbuch
Language
English
ISBN 10
3389179836
ISBN 13
9783389179833
Item weight
73 grams
Dimensions
210x148x4 mm

AHA-BUCH GmbH

Einbeck, Germany

5-star seller

AbeBooks seller since August 14, 2006

Shipping rates from Germany to U.S.A.

Item30 to 40 business days7 to 14 business days
First item£ 51.77£ 60.34
Delivery times are set by sellers and vary by carrier and location. Orders passing through Customs may face delays and buyers are responsible for any associated duties or fees. Sellers may contact you regarding additional charges to cover any increased costs to ship your items.

Payment methods

  • Visa
  • Mastercard
  • American Express
  • Apple Pay
  • Google Pay
  • Bank Wire Transfer
  • Check
  • Paypal

Store description

Das Unternehmen AHA-BUCH GmbH: Seit der Gründung von AHA-BUCH im Juli 2005 ist unser Hauptziel, zufriedenen Kunden so schnell und so preisgünstig wie möglich ihren Bücherwunsch zu erfüllen. Unsere Firma beschäftigt 16 Mitarbeiter, die nur ein Ziel kennen: den Kunden und seine Wünsche! Auf über 3700 m2 Fläche haben wir über 100.000 Bücher, Modernes Antiquariat und Spiele auf Lager.

Specialty

Kinderbücher & Kinderhör Casetten, German Books, Software, Natur & Tiere, Ratgeber, Sachbücher, Englische Bücher, Medizin & Gesundheit, Universität & Studium

Seller's business information

AHA-BUCH GmbH

Garlebsen 48
Einbeck, Germany 37574