15 December 2017

Google is using light beam tech to connect rural India to the internet


 Google is preparing to use light beams to bring rural areas of the planet online after it announced to a planned rollout in India. The firm is working with a telecom operator in Indian state Andhra Pradesh, home to over 50 million people, to use Free Space Optical Communications (FSOC), a technology that uses beams of light to deliver high-speed, high-capacity connectivity over long… Read More
Read Full Article

How to Unlock More Channels on Plex With the Unsupported App Store


plex-tricks-tips

Plex is most known for playing almost every video and audio file, regardless of format. It also includes fantastic metadata capabilities, meaning as long as you’ve correctly named the original media file, it can pull in episode synopses, cast lists, etc. But there’s another often underappreciated side to the app: TV Channels. Some channels are available in the official Channel Directory, but that’s only half the story. There are hundreds more channels available through the Unsupported App Store. Here’s how to install the Plex Unsupported App Store and unlock the extra content. How to Install the Plex Unsupported App Store Download...

Read the full article: How to Unlock More Channels on Plex With the Unsupported App Store


Read Full Article

How to Transcribe Audio to Text for Free


You might think you don’t have a need for a transcribing tool if you aren’t a journalist, lawyer, or medical professional, but you’d be surprised. You could be in a meeting with a recording app turned on. Or in a classroom lecture as your professor drones away. You can also be a writer, like me, who makes audio notes on the go. Basically, we can all benefit from ways to convert audio to text, transcription is the best way to do that. How to Transcribe Audio Using oTranscribe In the good old days, turning audio into text was sheer grunt work, but not...

Read the full article: How to Transcribe Audio to Text for Free


Read Full Article

The difference between good and bad Facebooking


 “Social media” is a clumsy term that entangles enriching social interaction with mindless media consumption. It’s a double-edged sword whose sides aren’t properly distinguished. Taken as a whole, we can’t decide if it “brings the world closer together” like Facebook’s new mission statement says, or leaves us depressed and isolated. Read More
Read Full Article

The difference between good and bad Facebooking


 “Social media” is a clumsy term that entangles enriching social interaction with mindless media consumption. It’s a double-edged sword whose sides aren’t properly distinguished. Taken as a whole, we can’t decide if it “brings the world closer together” like Facebook’s new mission statement says, or leaves us depressed and isolated. Read More

Read Full Article

You Can Now Unlock Pandora Premium by Watching Ads


Pandora boasts millions of users, but the vast majority of them listen to the free radio offering. Pandora Radio only lets you influence the songs being played by giving a thumbs-up or a thumbs-down. Which obviously isn’t as compelling as streaming services like Spotify and Apple Music. However, Pandora would like to encourage more free users to switch up to using Pandora Premium, and it has hit on a way of doing just that. From now on, people using Pandora Radio or Pandora Plus can unlock a free Pandora Premium session simply by watching a 15-second advert. Unlock a Pandora...

Read the full article: You Can Now Unlock Pandora Premium by Watching Ads


Read Full Article

Chromecast is back on Amazon


 Amazon will once again allow sales of Apple TV and Google Chromecast on its site, after banning them two years ago in an effort to promote its own Fire TV hardware to online shoppers. The addition of Apple TV is not surprising – Amazon and Apple came to an agreement earlier this year that included the launch of the Amazon Prime Video app for Apple TV, and the return of the Apple TV to… Read More
Read Full Article

Taylor Swift’s new app, The Swift Life, is out now


 In the midst of an reputational makeover, Taylor Swift is debuting a brand new app called The Swift Life. The app is a dedicated social network for Swift’s fans, letting them communicate with each other as well as get exclusive pictures, video, news and more direct from T-Swift herself. And they can communicate with Swift, too. Using an ‘extremely rare and valuable Taymoji&#8217… Read More
Read Full Article

Use the Waterfall Project Management Method to Organize Your Life


project-management-organize

The tools we typically use for managing personal projects have a major drawback. To-do lists and their variations (like Kanban boards) are fine to track what should be done and (maybe) who should do it. But they aren’t as good at planning out the when. This is the domain of business-oriented tools, such as the Waterfall project management approach. One of the primary reasons companies use a set of tools and processes to manage their projects is for profitability. Once management decides a particular project will add value to the business, there is a cost for every hour that project...

Read the full article: Use the Waterfall Project Management Method to Organize Your Life


Read Full Article

14 December 2017

Improving End-to-End Models For Speech Recognition




Traditional automatic speech recognition (ASR) systems, used for a variety of voice search applications at Google, are comprised of an acoustic model (AM), a pronunciation model (PM) and a language model (LM), all of which are independently trained, and often manually designed, on different datasets [1]. AMs take acoustic features and predict a set of subword units, typically context-dependent or context-independent phonemes. Next, a hand-designed lexicon (the PM) maps a sequence of phonemes produced by the acoustic model to words. Finally, the LM assigns probabilities to word sequences. Training independent components creates added complexities and is suboptimal compared to training all components jointly. Over the last several years, there has been a growing popularity in developing end-to-end systems, which attempt to learn these separate components jointly as a single system. While these end-to-end models have shown promising results in the literature [2, 3], it is not yet clear if such approaches can improve on current state-of-the-art conventional systems.

Today we are excited to share “State-of-the-art Speech Recognition With Sequence-to-Sequence Models [4],” which describes a new end-to-end model that surpasses the performance of a conventional production system [1]. We show that our end-to-end system achieves a word error rate (WER) of 5.6%, which corresponds to a 16% relative improvement over a strong conventional system which achieves a 6.7% WER. Additionally, the end-to-end model used to output the initial word hypothesis, before any hypothesis rescoring, is 18 times smaller than the conventional model, as it contains no separate LM and PM.

Our system builds on the Listen-Attend-Spell (LAS) end-to-end architecture, first presented in [2]. The LAS architecture consists of 3 components. The listener encoder component, which is similar to a standard AM, takes the a time-frequency representation of the input speech signal, x, and uses a set of neural network layers to map the input to a higher-level feature representation, henc. The output of the encoder is passed to an attender, which uses henc to learn an alignment between input features x and predicted subword units {yn, … y0}, where each subword is typically a grapheme or wordpiece. Finally, the output of the attention module is passed to the speller (i.e., decoder), similar to an LM, that produces a probability distribution over a set of hypothesized words.
Components of the LAS End-to-End Model.
All components of the LAS model are trained jointly as a single end-to-end neural network, instead of as separate modules like conventional systems, making it much simpler.
Additionally, because the LAS model is fully neural, there is no need for external, manually designed components such as finite state transducers, a lexicon, or text normalization modules. Finally, unlike conventional models, training end-to-end models does not require bootstrapping from decision trees or time alignments generated from a separate system, and can be trained given pairs of text transcripts and the corresponding acoustics.

In [4], we introduce a variety of novel structural improvements, including improving the attention vectors passed to the decoder and training with longer subword units (i.e., wordpieces). In addition, we also introduce numerous optimization improvements for training, including the use of minimum word error rate training [5]. These structural and optimization improvements are what accounts for obtaining the 16% relative improvement over the conventional model.

Another exciting potential application for this research is multi-dialect and multi-lingual systems, where the simplicity of optimizing a single neural network makes such a model very attractive. Here data for all dialects/languages can be combined to train one network, without the need for a separate AM, PM and LM for each dialect/language. We find that these models work well on 7 english dialects [6] and 8 Indian languages [7], while outperforming a model trained separately on each individual language/dialect.

While we are excited by our results, our work is not done. Currently, these models cannot process speech in real time [8, 9, 10], which is a strong requirement for latency-sensitive applications such as voice search. In addition, these models still compare negatively to production when evaluated on live production data. Furthermore, our end-to-end model is learned on 22,000 audio-text pair utterances compared to a conventional system that is typically trained on significantly larger corpora. In addition, our proposed model is not able to learn proper spellings for rarely used words such as proper nouns, which is normally performed with a hand-designed PM. Our ongoing efforts are focused now on addressing these challenges.

Acknowledgements
This work was done as a strong collaborative effort between Google Brain and Speech teams. Contributors include Tara Sainath, Rohit Prabhavalkar, Bo Li, Kanishka Rao, Shankar Kumar, Shubham Toshniwal, Michiel Bacchiani and Johan Schalkwyk from the Speech team; as well as Yonghui Wu, Patrick Nguyen, Zhifeng Chen, Chung-cheng Chiu, Anjuli Kannan, Ron Weiss and Navdeep Jaitly from the Google Brain team. The work is described in more detail in papers [4-11]

References
[1] G. Pundak and T. N. Sainath, “Lower Frame Rate Neural Network Acoustic Models," in Proc. Interspeech, 2016.

[2] W. Chan, N. Jaitly, Q. V. Le, and O. Vinyals, “Listen, attend and spell,” CoRR, vol. abs/1508.01211, 2015

[3] R. Prabhavalkar, K. Rao, T. N. Sainath, B. Li, L. Johnson, and N. Jaitly, “A Comparison of Sequence-to-sequence Models for Speech Recognition,” in Proc. Interspeech, 2017.

[4] C.C. Chiu, T.N. Sainath, Y. Wu, R. Prabhavalkar, P. Nguyen, Z. Chen, A. Kannan, R.J. Weiss, K. Rao, K. Gonina, N. Jaitly, B. Li, J. Chorowski and M. Bacchiani, “State-of-the-art Speech Recognition With Sequence-to-Sequence Models,” submitted to ICASSP 2018.

[5] R. Prabhavalkar, T.N. Sainath, Y. Wu, P. Nguyen, Z. Chen, C.C. Chiu and A. Kannan, “Minimum Word Error Rate Training for Attention-based Sequence-to-Sequence Models,” submitted to ICASSP 2018.

[6] B. Li, T.N. Sainath, K. Sim, M. Bacchiani, E. Weinstein, P. Nguyen, Z. Chen, Y. Wu and K. Rao, “Multi-Dialect Speech Recognition With a Single Sequence-to-Sequence Model” submitted to ICASSP 2018.

[7] S. Toshniwal, T.N. Sainath, R.J. Weiss, B. Li, P. Moreno, E. Weinstein and K. Rao, “End-to-End Multilingual Speech Recognition using Encoder-Decoder Models”, submitted to ICASSP 2018.

[8] T.N. Sainath, C.C. Chiu, R. Prabhavalkar, A. Kannan, Y. Wu, P. Nguyen and Z. Chen, “Improving the Performance of Online Neural Transducer Models”, submitted to ICASSP 2018.

[9] D. Lawson*, C.C. Chiu*, G. Tucker*, C. Raffel, K. Swersky, N. Jaitly. “Learning Hard Alignments with Variational Inference”, submitted to ICASSP 2018.

[10] T.N. Sainath, R. Prabhavalkar, S. Kumar, S. Lee, A. Kannan, D. Rybach, V. Schogol, P. Nguyen, B. Li, Y. Wu, Z. Chen and C.C. Chiu, “No Need for a Lexicon? Evaluating the Value of the Pronunciation Lexica in End-to-End Models,” submitted to ICASSP 2018.

[11] A. Kannan, Y. Wu, P. Nguyen, T.N. Sainath, Z. Chen and R. Prabhavalkar. “An Analysis of Incorporating an External Language Model into a Sequence-to-Sequence Model,” submitted to ICASSP 2018.

Facebook pushes pre-roll ads on Watch as it stops subsidizing Live


 Everyone’s least favorite ads are coming to Facebook, but six second pre-rolls will only appear on original Watch tab videos you purposefully view and not in the News Feed. Facebook is embracing pre-rolls after years of shunning them as it tries to make pay outs to video creators sustainable. Facebook’s head of video Fidji Simo tells TechCrunch that it will not renew direct… Read More
Read Full Article

Best Amazon Alexa Voice Commands for Phillips Hue


The Philips Hue personal wireless lighting system is an excellent way to make your dumb lightbulbs smart, but wouldn’t it be cool if you could talk to your lights? You know, say something like “Howdy there, if it’s not too much trouble, could you turn the lights on?” Well, thanks to Amazon Echo and Alexa, you can! Alexia Rolls Artsn Itln 3Chs Srd 10.5 Oz Alexia Rolls Artsn Itln 3Chs Srd 10.5 Oz Buy Now At Amazon Today I’ll be showing you the best voice commands for Hue and Alexa, although many of these commands will also work with Google...

Read the full article: Best Amazon Alexa Voice Commands for Phillips Hue


Read Full Article

A new version of Mixer, Microsoft’s Twitch rival, hits iOS and Android


 Microsoft today is officially launching a new version of its Mixer mobile gameplay streaming app, its Twitch rival. The app, which is initially available on Android with iOS arriving soon, was first introduced into beta testing this fall, with a focus on improvements to its overall user experience, content discovery, performance and personalization features. For example, the beta build… Read More

Read Full Article

Facebook pushes pre-roll ads on Watch as it stops subsidizing Live


 Everyone’s least favorite ads are coming to Facebook, but six second pre-rolls will only appear on original Watch tab videos you purposefully view and not in the News Feed. Facebook is embracing pre-rolls after years of shunning them as it tries to make pay outs to video creators sustainable. Facebook’s head of video Fidji Simo tells TechCrunch that it will not renew direct… Read More

Read Full Article

Google adds price tracking and deals to Google Flights, Google Trips and hotel search


 Google today is expanding its booking features for travelers using Google services including Trips, Flights and hotel search, with a focus on helping people find better rates. For example, Google can now tell you when’s the best time to buy an airline ticket or see when room rates are higher, among other things. These price-tracking features are similar to those that some other travel… Read More
Read Full Article

How to Dictate Email in Microsoft Outlook


dictate-email-outlook

It’s time to speed up how you write emails. If you struggle typing quickly, dictating email can help boost your productivity. We’re going to show you Dictate, which integrates straight into Outlook. Dictate is a utility developed by Microsoft and also works with other Office programs. You just plug your microphone in, click a button, and start talking. Everything you say is then transcribed. If you use Dictate or have a different speech-to-text software that you use, please let us know in the comments. About Dictate Microsoft Garage is a division of Microsoft that allows employees to work on their own projects...

Read the full article: How to Dictate Email in Microsoft Outlook


Read Full Article

10 Environmental Games That Teach Kids About Earth, Ecology & Conservation


The kids of today will inherit the Earth of tomorrow. They will also be left to clean up the mess we leave behind today. With the right ecological education, we can hope they will hit the ground running. A lot of schools and educational institutions are doing their bit by including the environment as part of the curriculum. Words like “carbon footprint” and “global warming” come to them as easily as the name of any present-day rock star. Technology has a trick. Environmental education can be taken out of musty textbooks and turned into interactive games. Like any other strategy game, kids...

Read the full article: 10 Environmental Games That Teach Kids About Earth, Ecology & Conservation


Read Full Article

Snapchat launches augmented reality developer platform Lens Studio


 Snapchat is finally opening up so outside developers can help it offer infinite augmented reality experiences beyond those it designs in-house. Today Snap launches the Lens Studio AR developer tool for desktops so anyone can create World Lenses that place interactive, imaginary 3D objects in your photos and videos. But brands, news publishers, and developers will have to promote their own… Read More
Read Full Article

Snapchat launches augmented reality developer platform Lens Studio


 Snapchat is finally opening up so outside developers can help it offer infinite augmented reality experiences beyond those it designs in-house. Today Snap launches the Lens Studio AR developer tool for desktops so anyone can create World Lenses that place interactive, imaginary 3D objects in your photos and videos. But brands, news publishers, and developers will have to promote their own… Read More

Read Full Article

A week on the wrist with the Alpina Startimer


 It’s refreshing to wear a mechanical watch. The soft sweep of the seconds hand reminds us of the fleeting nature of time while the endless ticking in a dark room is a comfort and a spur to action. Add in a little limited edition provenance with big face and crown and you’ve got a stew going. This particular stew is called the Alpina Startimer. It is a pilot’s watch, a watch… Read More

Read Full Article