Read Full Article
A blog about how-to, internet, social-networks, windows, linux, blogging, tips and tricks.
15 December 2017
Google is using light beam tech to connect rural India to the internet
Read Full Article
How to Unlock More Channels on Plex With the Unsupported App Store
Plex is most known for playing almost every video and audio file, regardless of format. It also includes fantastic metadata capabilities, meaning as long as you’ve correctly named the original media file, it can pull in episode synopses, cast lists, etc. But there’s another often underappreciated side to the app: TV Channels. Some channels are available in the official Channel Directory, but that’s only half the story. There are hundreds more channels available through the Unsupported App Store. Here’s how to install the Plex Unsupported App Store and unlock the extra content. How to Install the Plex Unsupported App Store Download...
Read the full article: How to Unlock More Channels on Plex With the Unsupported App Store
Read Full Article
How to Transcribe Audio to Text for Free
You might think you don’t have a need for a transcribing tool if you aren’t a journalist, lawyer, or medical professional, but you’d be surprised. You could be in a meeting with a recording app turned on. Or in a classroom lecture as your professor drones away. You can also be a writer, like me, who makes audio notes on the go. Basically, we can all benefit from ways to convert audio to text, transcription is the best way to do that. How to Transcribe Audio Using oTranscribe In the good old days, turning audio into text was sheer grunt work, but not...
Read the full article: How to Transcribe Audio to Text for Free
Read Full Article
The difference between good and bad Facebooking
Read Full Article
The difference between good and bad Facebooking
Read Full Article
You Can Now Unlock Pandora Premium by Watching Ads
Pandora boasts millions of users, but the vast majority of them listen to the free radio offering. Pandora Radio only lets you influence the songs being played by giving a thumbs-up or a thumbs-down. Which obviously isn’t as compelling as streaming services like Spotify and Apple Music. However, Pandora would like to encourage more free users to switch up to using Pandora Premium, and it has hit on a way of doing just that. From now on, people using Pandora Radio or Pandora Plus can unlock a free Pandora Premium session simply by watching a 15-second advert. Unlock a Pandora...
Read the full article: You Can Now Unlock Pandora Premium by Watching Ads
Read Full Article
Chromecast is back on Amazon
Read Full Article
Taylor Swift’s new app, The Swift Life, is out now
Read Full Article
Use the Waterfall Project Management Method to Organize Your Life
The tools we typically use for managing personal projects have a major drawback. To-do lists and their variations (like Kanban boards) are fine to track what should be done and (maybe) who should do it. But they aren’t as good at planning out the when. This is the domain of business-oriented tools, such as the Waterfall project management approach. One of the primary reasons companies use a set of tools and processes to manage their projects is for profitability. Once management decides a particular project will add value to the business, there is a cost for every hour that project...
Read the full article: Use the Waterfall Project Management Method to Organize Your Life
Read Full Article
14 December 2017
Improving End-to-End Models For Speech Recognition
Posted by Tara N. Sainath, Research Scientist, Speech Team and Yonghui Wu, Research Scientist, Google Brain Team
Traditional automatic speech recognition (ASR) systems, used for a variety of voice search applications at Google, are comprised of an acoustic model (AM), a pronunciation model (PM) and a language model (LM), all of which are independently trained, and often manually designed, on different datasets [1]. AMs take acoustic features and predict a set of subword units, typically context-dependent or context-independent phonemes. Next, a hand-designed lexicon (the PM) maps a sequence of phonemes produced by the acoustic model to words. Finally, the LM assigns probabilities to word sequences. Training independent components creates added complexities and is suboptimal compared to training all components jointly. Over the last several years, there has been a growing popularity in developing end-to-end systems, which attempt to learn these separate components jointly as a single system. While these end-to-end models have shown promising results in the literature [2, 3], it is not yet clear if such approaches can improve on current state-of-the-art conventional systems.
Today we are excited to share “State-of-the-art Speech Recognition With Sequence-to-Sequence Models [4],” which describes a new end-to-end model that surpasses the performance of a conventional production system [1]. We show that our end-to-end system achieves a word error rate (WER) of 5.6%, which corresponds to a 16% relative improvement over a strong conventional system which achieves a 6.7% WER. Additionally, the end-to-end model used to output the initial word hypothesis, before any hypothesis rescoring, is 18 times smaller than the conventional model, as it contains no separate LM and PM.
Our system builds on the Listen-Attend-Spell (LAS) end-to-end architecture, first presented in [2]. The LAS architecture consists of 3 components. The listener encoder component, which is similar to a standard AM, takes the a time-frequency representation of the input speech signal, x, and uses a set of neural network layers to map the input to a higher-level feature representation, henc. The output of the encoder is passed to an attender, which uses henc to learn an alignment between input features x and predicted subword units {yn, … y0}, where each subword is typically a grapheme or wordpiece. Finally, the output of the attention module is passed to the speller (i.e., decoder), similar to an LM, that produces a probability distribution over a set of hypothesized words.
| Components of the LAS End-to-End Model. |
Additionally, because the LAS model is fully neural, there is no need for external, manually designed components such as finite state transducers, a lexicon, or text normalization modules. Finally, unlike conventional models, training end-to-end models does not require bootstrapping from decision trees or time alignments generated from a separate system, and can be trained given pairs of text transcripts and the corresponding acoustics.
In [4], we introduce a variety of novel structural improvements, including improving the attention vectors passed to the decoder and training with longer subword units (i.e., wordpieces). In addition, we also introduce numerous optimization improvements for training, including the use of minimum word error rate training [5]. These structural and optimization improvements are what accounts for obtaining the 16% relative improvement over the conventional model.
Another exciting potential application for this research is multi-dialect and multi-lingual systems, where the simplicity of optimizing a single neural network makes such a model very attractive. Here data for all dialects/languages can be combined to train one network, without the need for a separate AM, PM and LM for each dialect/language. We find that these models work well on 7 english dialects [6] and 8 Indian languages [7], while outperforming a model trained separately on each individual language/dialect.
While we are excited by our results, our work is not done. Currently, these models cannot process speech in real time [8, 9, 10], which is a strong requirement for latency-sensitive applications such as voice search. In addition, these models still compare negatively to production when evaluated on live production data. Furthermore, our end-to-end model is learned on 22,000 audio-text pair utterances compared to a conventional system that is typically trained on significantly larger corpora. In addition, our proposed model is not able to learn proper spellings for rarely used words such as proper nouns, which is normally performed with a hand-designed PM. Our ongoing efforts are focused now on addressing these challenges.
Acknowledgements
This work was done as a strong collaborative effort between Google Brain and Speech teams. Contributors include Tara Sainath, Rohit Prabhavalkar, Bo Li, Kanishka Rao, Shankar Kumar, Shubham Toshniwal, Michiel Bacchiani and Johan Schalkwyk from the Speech team; as well as Yonghui Wu, Patrick Nguyen, Zhifeng Chen, Chung-cheng Chiu, Anjuli Kannan, Ron Weiss and Navdeep Jaitly from the Google Brain team. The work is described in more detail in papers [4-11]
References
[1] G. Pundak and T. N. Sainath, “Lower Frame Rate Neural Network Acoustic Models," in Proc. Interspeech, 2016.
[2] W. Chan, N. Jaitly, Q. V. Le, and O. Vinyals, “Listen, attend and spell,” CoRR, vol. abs/1508.01211, 2015
[3] R. Prabhavalkar, K. Rao, T. N. Sainath, B. Li, L. Johnson, and N. Jaitly, “A Comparison of Sequence-to-sequence Models for Speech Recognition,” in Proc. Interspeech, 2017.
[4] C.C. Chiu, T.N. Sainath, Y. Wu, R. Prabhavalkar, P. Nguyen, Z. Chen, A. Kannan, R.J. Weiss, K. Rao, K. Gonina, N. Jaitly, B. Li, J. Chorowski and M. Bacchiani, “State-of-the-art Speech Recognition With Sequence-to-Sequence Models,” submitted to ICASSP 2018.
[5] R. Prabhavalkar, T.N. Sainath, Y. Wu, P. Nguyen, Z. Chen, C.C. Chiu and A. Kannan, “Minimum Word Error Rate Training for Attention-based Sequence-to-Sequence Models,” submitted to ICASSP 2018.
[6] B. Li, T.N. Sainath, K. Sim, M. Bacchiani, E. Weinstein, P. Nguyen, Z. Chen, Y. Wu and K. Rao, “Multi-Dialect Speech Recognition With a Single Sequence-to-Sequence Model” submitted to ICASSP 2018.
[7] S. Toshniwal, T.N. Sainath, R.J. Weiss, B. Li, P. Moreno, E. Weinstein and K. Rao, “End-to-End Multilingual Speech Recognition using Encoder-Decoder Models”, submitted to ICASSP 2018.
[8] T.N. Sainath, C.C. Chiu, R. Prabhavalkar, A. Kannan, Y. Wu, P. Nguyen and Z. Chen, “Improving the Performance of Online Neural Transducer Models”, submitted to ICASSP 2018.
[9] D. Lawson*, C.C. Chiu*, G. Tucker*, C. Raffel, K. Swersky, N. Jaitly. “Learning Hard Alignments with Variational Inference”, submitted to ICASSP 2018.
[10] T.N. Sainath, R. Prabhavalkar, S. Kumar, S. Lee, A. Kannan, D. Rybach, V. Schogol, P. Nguyen, B. Li, Y. Wu, Z. Chen and C.C. Chiu, “No Need for a Lexicon? Evaluating the Value of the Pronunciation Lexica in End-to-End Models,” submitted to ICASSP 2018.
[11] A. Kannan, Y. Wu, P. Nguyen, T.N. Sainath, Z. Chen and R. Prabhavalkar. “An Analysis of Incorporating an External Language Model into a Sequence-to-Sequence Model,” submitted to ICASSP 2018.
Facebook pushes pre-roll ads on Watch as it stops subsidizing Live
Read Full Article
Best Amazon Alexa Voice Commands for Phillips Hue
The Philips Hue personal wireless lighting system is an excellent way to make your dumb lightbulbs smart, but wouldn’t it be cool if you could talk to your lights? You know, say something like “Howdy there, if it’s not too much trouble, could you turn the lights on?” Well, thanks to Amazon Echo and Alexa, you can! Alexia Rolls Artsn Itln 3Chs Srd 10.5 Oz Alexia Rolls Artsn Itln 3Chs Srd 10.5 Oz Buy Now At Amazon Today I’ll be showing you the best voice commands for Hue and Alexa, although many of these commands will also work with Google...
Read the full article: Best Amazon Alexa Voice Commands for Phillips Hue
Read Full Article
A new version of Mixer, Microsoft’s Twitch rival, hits iOS and Android
Read Full Article
Facebook pushes pre-roll ads on Watch as it stops subsidizing Live
Read Full Article
Google adds price tracking and deals to Google Flights, Google Trips and hotel search
Read Full Article
How to Dictate Email in Microsoft Outlook
It’s time to speed up how you write emails. If you struggle typing quickly, dictating email can help boost your productivity. We’re going to show you Dictate, which integrates straight into Outlook. Dictate is a utility developed by Microsoft and also works with other Office programs. You just plug your microphone in, click a button, and start talking. Everything you say is then transcribed. If you use Dictate or have a different speech-to-text software that you use, please let us know in the comments. About Dictate Microsoft Garage is a division of Microsoft that allows employees to work on their own projects...
Read the full article: How to Dictate Email in Microsoft Outlook
Read Full Article
10 Environmental Games That Teach Kids About Earth, Ecology & Conservation
The kids of today will inherit the Earth of tomorrow. They will also be left to clean up the mess we leave behind today. With the right ecological education, we can hope they will hit the ground running. A lot of schools and educational institutions are doing their bit by including the environment as part of the curriculum. Words like “carbon footprint” and “global warming” come to them as easily as the name of any present-day rock star. Technology has a trick. Environmental education can be taken out of musty textbooks and turned into interactive games. Like any other strategy game, kids...
Read the full article: 10 Environmental Games That Teach Kids About Earth, Ecology & Conservation
Read Full Article
Snapchat launches augmented reality developer platform Lens Studio
Read Full Article
Snapchat launches augmented reality developer platform Lens Studio
Read Full Article
A week on the wrist with the Alpina Startimer
Read Full Article