Web Crawling and Data Mining with Apache Nutch

Note moyenne 2,89
( 9 avis fournis par GoodReads )
 
9781783286850: Web Crawling and Data Mining with Apache Nutch
Présentation de l'éditeur :

Perform web crawling and apply data mining in your application

Overview

  • Learn to run your application on single as well as multiple machines
  • Customize search in your application as per your requirements
  • Acquaint yourself with storing crawled webpages in a database and use them according to your needs

In Detail

Apache Nutch helps you to create your own search engine and customize it according to your needs. You can integrate Apache Nutch very easily with your existing application and get the maximum benefit from it. It can be easily integrated with different components like Apache Hadoop, Eclipse, and MySQL.

"Web Crawling and Data Mining with Apache Nutch" shows you all the necessary steps to help you in crawling webpages for your application and using them to make your application searching more efficient. You will create your own search engine and will be able to improve your application page rank in searching.

"Web Crawling and Data Mining with Apache Nutch" starts with the basics of crawling webpages for your application. You will learn to deploy Apache Solr on server containing data crawled by Apache Nutch and perform Sharding with Apache Nutch using Apache Solr.

You will integrate your application with databases such as MySQL, Hbase, and Accumulo, and also with Apache Solr, which is used as a searcher.

With this book, you will gain the necessary skills to create your own search engine. You will also perform link analysis and scoring that are helpful in improving the rank of your application page.

What you will learn from this book

  • Carry out web crawling for your application
  • Make your application searching efficient by integrating it with Apache Solr
  • Integrate your application with different databases for data storage purposes
  • Run your application in a cluster environment by integrating it with Apache Hadoop
  • Perform crawling operations with Eclipse, which is used as an IDE instead of the command line
  • Create your own plugin in Apache Nutch
  • Integrate Apache Solr with Apache Nutch, and deploy Apache Solr on Apache Tomcat
  • Apply Sharding on Apache Tomcat for getting good results from Apache Solr while searching

Approach

This book is a user-friendly guide that covers all the necessary steps and examples related to web crawling and data mining using Apache Nutch.

Who this book is written for

"Web Crawling and Data Mining with Apache Nutch" is aimed at data analysts, application developers, web mining engineers, and data scientists. It is a good start for those who want to learn how web crawling and data mining is applied in the current business world. It would be an added benefit for those who have some knowledge of web crawling and data mining.

Biographie de l'auteur :

Dr. Zakir Laliwala

Dr. Zakir Laliwala is an entrepreneur, an open source specialist, and a hands-on CTO at Attune Infocom. Attune Infocom provides enterprise open source solutions and services for SOA, BPM, ESB, Portal, cloud computing, and ECM. At Attune Infocom, he is responsible for product development and the delivery of solutions and services. He explores new enterprise open source technologies and defines architecture, roadmaps, and best practices. He has provided consultations and training to corporations around the world on various open source technologies such as Mule ESB, Activiti BPM, JBoss jBPM and Drools, Liferay Portal, Alfresco ECM, JBoss SOA, and cloud computing.

He received a Ph.D. in Information and Communication Technology from Dhirubhai Ambani Institute of Information and Communication Technology. He was an adjunct faculty at Dhirubhai Ambani Institute of Information and Communication Technology (DA-IICT), and he taught Master's degree students at CEPT.

He has published many research papers on web services, SOA, grid computing, and the semantic web in IEEE, and has participated in ACM International Conferences. He serves as a reviewer at various international conferences and journals. He has also published book chapters and written books on open source technologies. He was a co-author of the books Mule ESB Cookbook and Activiti5 Business Process Management Beginner's Guide, Packt Publishing.



Abdulbasit Shaikh

Abdulbasit Shaikh has more than two years of experience in the IT industry. He completed his Masters' degree from the Dhirubhai Ambani Institute of Information and Communication Technology (DA-IICT). He has a lot of experience in open source technologies. He has worked on a number of open source technologies, such as Apache Hadoop, Apache Solr, Apache ZooKeeper, Apache Mahout, Apache Nutch, and Liferay. He has provided training on Apache Nutch, Apache Hadoop, Apache Mahout, and AWS architect. He is currently working on the OpenStack technology. He has also delivered projects and training on open source technologies. He has a very good knowledge of cloud computing, such as AWS and Microsoft Azure, as he has successfully delivered many projects in cloud computing.

He is a very enthusiastic and active person when he is working on a project or delivering a project. Currently, he is working as a Java developer at Attune Infocom Pvt. Ltd. He is totally focused on open source technologies, and he is very much interested in sharing his knowledge with the open source community.

Les informations fournies dans la section « A propos du livre » peuvent faire référence à une autre édition de ce titre.

Meilleurs résultats de recherche sur AbeBooks

1.

Dr. Zakir Laliwala, Abdulbasit. Fazalmehmod Shaikh
Edité par Packt Publishing Limited, United Kingdom (2013)
ISBN 10 : 1783286857 ISBN 13 : 9781783286850
Neuf(s) Paperback Quantité : 10
impression à la demande
Vendeur
The Book Depository
(London, Royaume-Uni)
Evaluation vendeur
[?]

Description du livre Packt Publishing Limited, United Kingdom, 2013. Paperback. État : New. 234 x 190 mm. Language: English . Brand New Book ***** Print on Demand *****.Apache Nutch helps you to create your own search engine and customize it according to your needs. You can integrate Apache Nutch very easily with your existing application and get the maximum benefit from it. It can be easily integrated with different components like Apache Hadoop, Eclipse, and MySQL. Web Crawling and Data Mining with Apache Nutch shows you all the necessary steps to help you in crawling webpages for your application and using them to make your application searching more efficient. You will create your own search engine and will be able to improve your application page rank in searching. Web Crawling and Data Mining with Apache Nutch starts with the basics of crawling webpages for your application. You will learn to deploy Apache Solr on server containing data crawled by Apache Nutch and perform Sharding with Apache Nutch using Apache Solr. You will integrate your application with databases such as MySQL, Hbase, and Accumulo, and also with Apache Solr, which is used as a searcher. With this book, you will gain the necessary skills to create your own search engine. You will also perform link analysis and scoring that are helpful in improving the rank of your application page. N° de réf. du libraire AAV9781783286850

Plus d'informations sur ce vendeur | Poser une question au libraire

Acheter neuf
EUR 30
Autre devise

Ajouter au panier

Frais de port : Gratuit
De Royaume-Uni vers Etats-Unis
Destinations, frais et délais

2.

Laliwala, Zakir
Edité par Packt Publishing (2016)
ISBN 10 : 1783286857 ISBN 13 : 9781783286850
Neuf(s) Paperback Quantité : 1
impression à la demande
Vendeur
Ria Christie Collections
(Uxbridge, Royaume-Uni)
Evaluation vendeur
[?]

Description du livre Packt Publishing, 2016. Paperback. État : New. PRINT ON DEMAND Book; New; Publication Year 2016; Not Signed; Fast Shipping from the UK. No. book. N° de réf. du libraire ria9781783286850_lsuk

Plus d'informations sur ce vendeur | Poser une question au libraire

Acheter neuf
EUR 31,29
Autre devise

Ajouter au panier

Frais de port : EUR 3,88
De Royaume-Uni vers Etats-Unis
Destinations, frais et délais

3.

Dr. Zakir Laliwala, Abdulbasit. Fazalmehmod Shaikh
Edité par Packt Publishing Limited, United Kingdom (2013)
ISBN 10 : 1783286857 ISBN 13 : 9781783286850
Neuf(s) Paperback Quantité : 10
impression à la demande
Vendeur
The Book Depository US
(London, Royaume-Uni)
Evaluation vendeur
[?]

Description du livre Packt Publishing Limited, United Kingdom, 2013. Paperback. État : New. 234 x 190 mm. Language: English . Brand New Book ***** Print on Demand *****. Apache Nutch helps you to create your own search engine and customize it according to your needs. You can integrate Apache Nutch very easily with your existing application and get the maximum benefit from it. It can be easily integrated with different components like Apache Hadoop, Eclipse, and MySQL. Web Crawling and Data Mining with Apache Nutch shows you all the necessary steps to help you in crawling webpages for your application and using them to make your application searching more efficient. You will create your own search engine and will be able to improve your application page rank in searching. Web Crawling and Data Mining with Apache Nutch starts with the basics of crawling webpages for your application. You will learn to deploy Apache Solr on server containing data crawled by Apache Nutch and perform Sharding with Apache Nutch using Apache Solr. You will integrate your application with databases such as MySQL, Hbase, and Accumulo, and also with Apache Solr, which is used as a searcher. With this book, you will gain the necessary skills to create your own search engine. You will also perform link analysis and scoring that are helpful in improving the rank of your application page. N° de réf. du libraire AAV9781783286850

Plus d'informations sur ce vendeur | Poser une question au libraire

Acheter neuf
EUR 36,10
Autre devise

Ajouter au panier

Frais de port : Gratuit
De Royaume-Uni vers Etats-Unis
Destinations, frais et délais

4.

Dr. Zakir Laliwala
Edité par Packt Publishing Limited (2013)
ISBN 10 : 1783286857 ISBN 13 : 9781783286850
Neuf(s) Quantité : > 20
impression à la demande
Vendeur
Books2Anywhere
(Fairford, GLOS, Royaume-Uni)
Evaluation vendeur
[?]

Description du livre Packt Publishing Limited, 2013. PAP. État : New. New Book. Delivered from our UK warehouse in 3 to 5 business days. THIS BOOK IS PRINTED ON DEMAND. Established seller since 2000. N° de réf. du libraire LQ-9781783286850

Plus d'informations sur ce vendeur | Poser une question au libraire

Acheter neuf
EUR 26,67
Autre devise

Ajouter au panier

Frais de port : EUR 10,45
De Royaume-Uni vers Etats-Unis
Destinations, frais et délais

5.

Dr. Zakir Laliwala
Edité par Packt Publishing Limited (2013)
ISBN 10 : 1783286857 ISBN 13 : 9781783286850
Neuf(s) Quantité : > 20
impression à la demande
Vendeur
PBShop
(Wood Dale, IL, Etats-Unis)
Evaluation vendeur
[?]

Description du livre Packt Publishing Limited, 2013. PAP. État : New. New Book. Shipped from US within 10 to 14 business days. THIS BOOK IS PRINTED ON DEMAND. Established seller since 2000. N° de réf. du libraire IQ-9781783286850

Plus d'informations sur ce vendeur | Poser une question au libraire

Acheter neuf
EUR 33,52
Autre devise

Ajouter au panier

Frais de port : EUR 3,70
Vers Etats-Unis
Destinations, frais et délais

6.

Dr Zakir Laliwala,Abdulbasit Fazalmehmod Shaikh,Za
Edité par Packt Publishing (2017)
ISBN 10 : 1783286857 ISBN 13 : 9781783286850
Neuf(s) Paperback Quantité : 20
impression à la demande
Vendeur
Murray Media
(North Miami Beach, FL, Etats-Unis)
Evaluation vendeur
[?]

Description du livre Packt Publishing, 2017. Paperback. État : New. This item is printed on demand. N° de réf. du libraire 1783286857

Plus d'informations sur ce vendeur | Poser une question au libraire

Acheter neuf
EUR 42,37
Autre devise

Ajouter au panier

Frais de port : EUR 2,77
Vers Etats-Unis
Destinations, frais et délais

7.

Dr Zakir Laliwala,Abdulbasit Fazalmehmod Shaikh,Zakir Laliwala
Edité par Packt Publishing
ISBN 10 : 1783286857 ISBN 13 : 9781783286850
Neuf(s) Paperback Quantité : 1
Vendeur
Irish Booksellers
(Rumford, ME, Etats-Unis)
Evaluation vendeur
[?]

Description du livre Packt Publishing. Paperback. État : New. book. N° de réf. du libraire 1783286857

Plus d'informations sur ce vendeur | Poser une question au libraire

Acheter neuf
EUR 48,68
Autre devise

Ajouter au panier

Frais de port : Gratuit
Vers Etats-Unis
Destinations, frais et délais

8.

Dr Zakir Laliwala,Abdulbasit Fazalmehmod Shaikh,Zakir Laliwala
Edité par Packt Publishing
ISBN 10 : 1783286857 ISBN 13 : 9781783286850
Neuf(s) PAPERBACK Quantité : > 20
Vendeur
Russell Books
(Victoria, BC, Canada)
Evaluation vendeur
[?]

Description du livre Packt Publishing. PAPERBACK. État : New. 1783286857 Special order direct from the distributor. N° de réf. du libraire ING9781783286850

Plus d'informations sur ce vendeur | Poser une question au libraire

Acheter neuf
EUR 50,13
Autre devise

Ajouter au panier

Frais de port : EUR 6,49
De Canada vers Etats-Unis
Destinations, frais et délais

9.

Abdulbasit Fazalmehmod Shaikh, Zakir Laliwala Dr Zakir Laliwala
Edité par Packt Publishing
ISBN 10 : 1783286857 ISBN 13 : 9781783286850
Neuf(s) Paperback Quantité : 20
Vendeur
BuySomeBooks
(Las Vegas, NV, Etats-Unis)
Evaluation vendeur
[?]

Description du livre Packt Publishing. Paperback. État : New. Paperback. Dimensions: 9.2in. x 7.5in. x 0.4in. This item ships from multiple locations. Your book may arrive from Roseburg,OR, La Vergne,TN. Paperback. N° de réf. du libraire 9781783286850

Plus d'informations sur ce vendeur | Poser une question au libraire

Acheter neuf
EUR 53,77
Autre devise

Ajouter au panier

Frais de port : EUR 3,66
Vers Etats-Unis
Destinations, frais et délais

10.

Abdulbasit Shaikh
ISBN 10 : 1783286857 ISBN 13 : 9781783286850
Neuf(s) Couverture souple Quantité : 15
impression à la demande
Vendeur
English-Book-Service Mannheim
(Mannheim, Allemagne)
Evaluation vendeur
[?]

Description du livre État : New. This item is printed on demand for shipment within 3 working days. N° de réf. du libraire LP9781783286850

Plus d'informations sur ce vendeur | Poser une question au libraire

Acheter neuf
EUR 55,11
Autre devise

Ajouter au panier

Frais de port : EUR 5
De Allemagne vers Etats-Unis
Destinations, frais et délais

autres exemplaires de ce livre sont disponibles

Afficher tous les résultats pour ce livre