Click here to close now.


Python Authors: XebiaLabs Blog, Hovhannes Avoyan, Carmen Gonzalez, Ignacio M. Llorente, Elizabeth White

Blog Feed Post

Talend Simplifies Big Data Further with New Release of Enterprise Open Source Integration Platform

Open source software leader advances next-generation integration solution with big data profiling for Hadoop, support of major NoSQL databases and increased usability features

Maidenhead, UK - 5 November 2012 - Talend, a global open source software leader, today announced the availability of version 5.2 of its next-generation integration platform, the only offering that provides a unified environment for managing the entire lifecycle across data, application and process integration requirements. With version 5.2, Talend extends the industry's most flexible, scalable and adaptive integration platform with the introduction of key new capabilities, including big data profiling for Hadoop, support for widely-used and deployed NoSQL databases, and a set of improvements that increases product usability and performance across the entire platform.

Big Data Profiling
In its mission to democratise big data, Talend has focused extensively on solutions that make deploying and managing Apache Hadoop and related technologies simple, without requiring specific expertise in these areas. With version 5.2, Talend has taken its big data strategy a step further by adding big data profiling for Hadoop, providing companies with the ability to discover and understand data in Hadoop clusters. Among the typical problems associated with data quality are duplication, incompleteness and inconsistency, which create inefficiencies in data processing. Talend Platform for Big Data includes new capabilities for visibility into big data in all its forms and locations. These include the ability to analyse data in Hive databases on Hadoop "in place" without extraction and the ability to perform data hygiene tasks, including data cleansing, enrichment, matching and de-duplication directly inside the Hadoop cluster through Hadoop code generation.

Simplified NoSQL Integration with Hadoop
Talend 5.2 adds support for NoSQL databases in its integration solutions, Talend Platform for Big Data and Talend Open Studio for Big Data, with an initial set of connectors for Cassandra, HBase and MongoDB. Built on Talend's award-winning open source integration technology, Talend Open Studio for Big Data is a powerful and versatile open source solution for big data integration that natively supports Apache Hadoop, including connectors for Hadoop Distributed File System (HDFS), HCatalog, Hive, Oozie, Pig and Sqoop - in addition to the more than 450 connectors included natively in the product. As NoSQL has become the go-to technology for certain data architectures, the integration of these platforms into Talend's big data solution enables customers to use these new connectors to migrate and synchronise data between NoSQL databases and all other data stores and systems.

"Talend version 5.2 delivers on our vision of simplifying the development, integration and management of big data so that businesses can focus on using that data to make faster and more informed decisions," said Fabrice Bonan, co-founder and chief technical officer, Talend. "We provide the most powerful and versatile open source, big data solution to help organisations load, extract and improve disparate data while leveraging the massively parallel processing power of big data technologies including Apache Hadoop and leading NoSQL databases."

Latest Release of Talend's Integration Products
In addition to Talend's big data enhancements, Talend introduces version 5.2 of its flagship data integration products that leverage the Talend Unified Platform. New features focus on product usability, user productivity improvements and performance to provide a more robust and easier to use solution.

  • Talend Enterprise Data Integration - In v5.2, parallel execution of jobs can now leverage multi-core hardware. This new version also supports continuous integration between development, test and production environments and is integrated with open source build manager Maven.

  • Talend Enterprise Data Quality - Version 5.2 includes expanded address validation algorithms, precise e-mail validity detection, and native fraud detection capabilities. A new set of components allows customers to use Melissa Data to validate addresses.

  • Talend Enterprise MDM - Support for a wider range of enterprise architectures in v5.2 lowers the barrier to MDM adoption; organisations can now use their Oracle, MySQL, Derby or H2 databases as the underlying MDM data store.

  • Talend Enterprise ESB - In this version, Continuous Integration between development, test and production environments is now available. Version control system Nexus is also supported for versioning and deployment.

  • Talend Enterprise BPM - Talend v5.2 presents a fully integrated BPM engine into the Talend Runtime. Talend users only need to manage a single container, which can run data jobs, web services, REST applications and now BPM processes. With fewer moving parts in system environments and the flexibility to run multiple instances of different application types within the same container, the work of the IT administrator is significantly reduced in terms of management and maintenance of the software.

Version 5.2 of Talend Open Studio for Data Integration, Talend Open Studio for Data Quality, Talend Open Studio for MDM, Talend Open Studio for ESB and Talend Open Studio for Big Data are available for immediate download from Talend's web site Version 5.2 of the commercial subscription products, available before the end of 2012, will be provided to all existing Talend customers as part of their subscription agreement and can be procured through the usual Talend representatives or partners.

About Talend
Talend is the recognised market leader in open source integration solutions. The company's enterprise integration platform helps organisations minimise costs and maximise the value of data integration, ETL, data quality, master data management, application integration and business process management, while supporting their shift toward the Cloud and Big Data. More than 3,500 paying customers worldwide, including eBay, ING, The Weather Channel, Deutsche Post and Allianz, subscribe to Talend's solutions and services. With over 20 million downloads, Talend's products are the most trusted integration solutions in the world. The company has major offices in North America, Europe and Asia, and a global network of technical and services partners. For more information, please visit


PR Contacts:
Selene Regan
[email protected]

Tom Webb
01252 727313
[email protected]

Read the original blog entry...

More Stories By RealWire News Distribution

RealWire is a global news release distribution service specialising in the online media. The RealWire approach focuses on delivering relevant content to the receivers of our client's news releases. As we know that it is only through delivering relevance, that influence can ever be achieved.

@ThingsExpo Stories
Continuous processes around the development and deployment of applications are both impacted by -- and a benefit to -- the Internet of Things trend. To help better understand the relationship between DevOps and a plethora of new end-devices and data please welcome Gary Gruver, consultant, author and a former IT executive who has led many large-scale IT transformation projects, and John Jeremiah, Technology Evangelist at Hewlett Packard Enterprise (HPE), on Twitter at @j_jeremiah. The discussion is moderated by me, Dana Gardner, Principal Analyst at Interarbor Solutions.
Too often with compelling new technologies market participants become overly enamored with that attractiveness of the technology and neglect underlying business drivers. This tendency, what some call the “newest shiny object syndrome” is understandable given that virtually all of us are heavily engaged in technology. But it is also mistaken. Without concrete business cases driving its deployment, IoT, like many other technologies before it, will fade into obscurity.
With all the incredible momentum behind the Internet of Things (IoT) industry, it is easy to forget that not a single CEO wakes up and wonders if “my IoT is broken.” What they wonder is if they are making the right decisions to do all they can to increase revenue, decrease costs, and improve customer experience – effectively the same challenges they have always had in growing their business. The exciting thing about the IoT industry is now these decisions can be better, faster, and smarter. Now all corporate assets – people, objects, and spaces – can share information about themselves and thei...
The Internet of Things is clearly many things: data collection and analytics, wearables, Smart Grids and Smart Cities, the Industrial Internet, and more. Cool platforms like Arduino, Raspberry Pi, Intel's Galileo and Edison, and a diverse world of sensors are making the IoT a great toy box for developers in all these areas. In this Power Panel at @ThingsExpo, moderated by Conference Chair Roger Strukhoff, panelists discussed what things are the most important, which will have the most profound effect on the world, and what should we expect to see over the next couple of years.
Discussions of cloud computing have evolved in recent years from a focus on specific types of cloud, to a world of hybrid cloud, and to a world dominated by the APIs that make today's multi-cloud environments and hybrid clouds possible. In this Power Panel at 17th Cloud Expo, moderated by Conference Chair Roger Strukhoff, panelists addressed the importance of customers being able to use the specific technologies they need, through environments and ecosystems that expose their APIs to make true change and transformation possible.
The cloud. Like a comic book superhero, there seems to be no problem it can’t fix or cost it can’t slash. Yet making the transition is not always easy and production environments are still largely on premise. Taking some practical and sensible steps to reduce risk can also help provide a basis for a successful cloud transition. A plethora of surveys from the likes of IDG and Gartner show that more than 70 percent of enterprises have deployed at least one or more cloud application or workload. Yet a closer inspection at the data reveals less than half of these cloud projects involve production...
Microservices are a very exciting architectural approach that many organizations are looking to as a way to accelerate innovation. Microservices promise to allow teams to move away from monolithic "ball of mud" systems, but the reality is that, in the vast majority of organizations, different projects and technologies will continue to be developed at different speeds. How to handle the dependencies between these disparate systems with different iteration cycles? Consider the "canoncial problem" in this scenario: microservice A (releases daily) depends on a couple of additions to backend B (re...
Growth hacking is common for startups to make unheard-of progress in building their business. Career Hacks can help Geek Girls and those who support them (yes, that's you too, Dad!) to excel in this typically male-dominated world. Get ready to learn the facts: Is there a bias against women in the tech / developer communities? Why are women 50% of the workforce, but hold only 24% of the STEM or IT positions? Some beginnings of what to do about it! In her Day 2 Keynote at 17th Cloud Expo, Sandy Carter, IBM General Manager Cloud Ecosystem and Developers, and a Social Business Evangelist, wil...
PubNub has announced the release of BLOCKS, a set of customizable microservices that give developers a simple way to add code and deploy features for realtime apps.PubNub BLOCKS executes business logic directly on the data streaming through PubNub’s network without splitting it off to an intermediary server controlled by the customer. This revolutionary approach streamlines app development, reduces endpoint-to-endpoint latency, and allows apps to better leverage the enormous scalability of PubNub’s Data Stream Network.
Apps and devices shouldn't stop working when there's limited or no network connectivity. Learn how to bring data stored in a cloud database to the edge of the network (and back again) whenever an Internet connection is available. In his session at 17th Cloud Expo, Ben Perlmutter, a Sales Engineer with IBM Cloudant, demonstrated techniques for replicating cloud databases with devices in order to build offline-first mobile or Internet of Things (IoT) apps that can provide a better, faster user experience, both offline and online. The focus of this talk was on IBM Cloudant, Apache CouchDB, and ...
Container technology is shaping the future of DevOps and it’s also changing the way organizations think about application development. With the rise of mobile applications in the enterprise, businesses are abandoning year-long development cycles and embracing technologies that enable rapid development and continuous deployment of apps. In his session at DevOps Summit, Kurt Collins, Developer Evangelist at, examined how Docker has evolved into a highly effective tool for application delivery by allowing increasingly popular Mobile Backend-as-a-Service (mBaaS) platforms to quickly crea...
I recently attended and was a speaker at the 4th International Internet of @ThingsExpo at the Santa Clara Convention Center. I also had the opportunity to attend this event last year and I wrote a blog from that show talking about how the “Enterprise Impact of IoT” was a key theme of last year’s show. I was curious to see if the same theme would still resonate 365 days later and what, if any, changes I would see in the content presented.
Cloud computing delivers on-demand resources that provide businesses with flexibility and cost-savings. The challenge in moving workloads to the cloud has been the cost and complexity of ensuring the initial and ongoing security and regulatory (PCI, HIPAA, FFIEC) compliance across private and public clouds. Manual security compliance is slow, prone to human error, and represents over 50% of the cost of managing cloud applications. Determining how to automate cloud security compliance is critical to maintaining positive ROI. Raxak Protect is an automated security compliance SaaS platform and ma...
Internet of @ThingsExpo, taking place June 7-9, 2016 at Javits Center, New York City and Nov 1-3, 2016, at the Santa Clara Convention Center in Santa Clara, CA, is co-located with the 18th International @CloudExpo and will feature technical sessions from a rock star conference faculty and the leading industry players in the world and ThingsExpo New York Call for Papers is now open.
With major technology companies and startups seriously embracing IoT strategies, now is the perfect time to attend @ThingsExpo 2016 in New York and Silicon Valley. Learn what is going on, contribute to the discussions, and ensure that your enterprise is as "IoT-Ready" as it can be! Internet of @ThingsExpo, taking place Nov 3-5, 2015, at the Santa Clara Convention Center in Santa Clara, CA, is co-located with 17th Cloud Expo and will feature technical sessions from a rock star conference faculty and the leading industry players in the world. The Internet of Things (IoT) is the most profound cha...
We are rapidly moving to a brave new world of interconnected smart homes, cars, offices and factories known as the Internet of Things (IoT). Sensors and monitoring devices will touch every part of our lives. Let's take a closer look at the Internet of Things. The Internet of Things is a worldwide network of objects and devices connected to the Internet. They are electronics, sensors, software and more. These objects connect to the Internet and can be controlled remotely via apps and programs. Because they can be accessed via the Internet, these devices create a tremendous opportunity to inte...
Today air travel is a minefield of delays, hassles and customer disappointment. Airlines struggle to revitalize the experience. GE and M2Mi will demonstrate practical examples of how IoT solutions are helping airlines bring back personalization, reduce trip time and improve reliability. In their session at @ThingsExpo, Shyam Varan Nath, Principal Architect with GE, and Dr. Sarah Cooper, M2Mi’s VP Business Development and Engineering, explored the IoT cloud-based platform technologies driving this change including privacy controls, data transparency and integration of real time context with p...
We all know that data growth is exploding and storage budgets are shrinking. Instead of showing you charts on about how much data there is, in his General Session at 17th Cloud Expo, Scott Cleland, Senior Director of Product Marketing at HGST, showed how to capture all of your data in one place. After you have your data under control, you can then analyze it in one place, saving time and resources.
The Internet of Things (IoT) is growing rapidly by extending current technologies, products and networks. By 2020, Cisco estimates there will be 50 billion connected devices. Gartner has forecast revenues of over $300 billion, just to IoT suppliers. Now is the time to figure out how you’ll make money – not just create innovative products. With hundreds of new products and companies jumping into the IoT fray every month, there’s no shortage of innovation. Despite this, McKinsey/VisionMobile data shows "less than 10 percent of IoT developers are making enough to support a reasonably sized team....
Just over a week ago I received a long and loud sustained applause for a presentation I delivered at this year’s Cloud Expo in Santa Clara. I was extremely pleased with the turnout and had some very good conversations with many of the attendees. Over the next few days I had many more meaningful conversations and was not only happy with the results but also learned a few new things. Here is everything I learned in those three days distilled into three short points.