|
所在平台: Udemy |
课程主页: https://www.udemy.com/course/learn-apache-solr-with-big-data-and-cloud-computing/
课程评论:没有评论
课程名称:学习Apache Solr与大数据和云计算 课程概述:Apache Solr是一个流行、快速的开源企业搜索平台,源自Apache Lucene项目。Solr的主要特点包括强大的全文搜索、命中高亮、分面搜索、近实时索引、动态集群、数据库集成、丰富文档处理(如Word、PDF)和地理空间搜索。Solr的高可靠性、可扩展性和容错性,使其能够支持分布式索引、复制和负载均衡查询、自动故障恢复和集中配置等功能。Solr被许多全球最大的网站用于搜索和导航功能。Solr是用Java编写的,运行在像Jetty这样的Servlet容器中,使用Lucene Java搜索库进行全文索引和搜索,并拥有REST风格的HTTP/XML及JSON API,可以方便地与几乎任何编程语言配合使用。Solr的外部配置功能强大,可以根据几乎任何类型的应用进行定制,无需Java编码,还具备广泛的插件架构,以满足更高级的自定义需求。 课程内容包括: - Solr的独立企业搜索服务器与REST风格API; - 文档索引与查询的多种格式(XML、JSON、CSV等); - 提供高级的全文搜索能力,优化高流量的网页流量; - 支持动态字段、文本分析、重用和组合字段功能; - 配置和监控接口,包括基于JMX的服务器统计信息; - SolrCloud支持自动化分布式索引和近实时索引; - 详尽的文档解析与索引能力,包括PDF、Word、HTML等格式。 本课程为希望了解和应用Solr进行搜索服务构建的学习者提供了全面的知识体系,适合希望掌握大数据和云计算环境中搜索技术的开发者和数据科学家。
Solr is the popular, blazing fast open source enterprise search platform from the Apache LuceneTMproject. Its major features include powerful full-text search, hit highlighting, faceted search, near real-time indexing, dynamic clustering, database integration, rich document (e.g., Word, PDF) handling, and geospatial search. Solr is highly reliable, scalable and fault tolerant, providing distributed indexing, replication and load-balanced querying, automated failover and recovery, centralized configuration and more. Solr powers the search and navigation features of many of the world's largest internet sites. Solr is written in Java and runs as a standalone full-text search server within a servlet container such as Jetty. Solr uses the Lucene Java search library at its core for full-text indexing and search, and has REST-like HTTP/XML and JSON APIs that make it easy to use from virtually any programming language. Solr's powerful external configuration allows it to be tailored to almost any type of application without Java coding, and it has an extensive plugin architecture when more advanced customization is required. Solr Features Solr is a standalone enterprise search server with a REST-like API. You put documents in it (called "indexing") via XML, JSON, CSV or binary over HTTP. You query it via HTTP GET and receive XML, JSON, CSV or binary results. Advanced Full-Text Search Capabilities Optimized for High Volume Web Traffic Standards Based Open Interfaces - XML, JSON and HTTP Comprehensive HTML Administration Interfaces Server statistics exposed over JMX for monitoring Linearly scalable, auto index replication, auto failover and recovery Near Real-time indexing Flexible and Adaptable with XML configuration Extensible Plugin Architecture Solr Uses the LuceneTM Search Library and Extends it! A Real Data Schema, with Numeric Types, Dynamic Fields, Unique Keys Powerful Extensions to the Lucene Query Language Faceted Search and Filtering Geospatial Search with support for multiple points per document and geo polygons Advanced, Configurable Text Analysis Highly Configurable and User Extensible Caching Performance Optimizations External Configuration via XML An AJAX based administration interface Monitorable Logging Fast near real-time incremental indexing and index replication Highly Scalable Distributed search with sharded index across multiple hosts JSON, XML, CSV/delimited-text, and binary update formats Easy ways to pull in data from databases and XML files from local disk and HTTP sources Rich Document Parsing and Indexing (PDF, Word, HTML, etc) using Apache Tika Apache UIMA integration for configurable metadata extraction Multiple search indices Detailed Features Schema Defines the field types and fields of documents Can drive more intelligent processing Declarative Lucene Analyzer specification Dynamic Fields enables on-the-fly addition of new fields CopyField functionality allows indexing a single field multiple ways, or combining multiple fields into a single searchable field Explicit types eliminates the need for guessing types of fields External file-based configuration of stopword lists, synonym lists, and protected word lists Many additional text analysis components including word splitting, regex and sounds-like filters Pluggable similarity model per field Query HTTP interface with configurable response formats (XML/XSLT, JSON, Python, Ruby, PHP, Velocity, CSV, binary) Sort by any number of fields, and by complex functions of numeric fields Advanced DisMax query parser for high relevancy results from user-entered queries Highlighted context snippets Faceted Searching based on unique field values, explicit queries, date ranges, numeric ranges or pivot Multi-Select Faceting by tagging and selectively excluding filters Spelling suggestions for user queries More Like This suggestions for given document Function Query - influence the score by user specified complex functions of numeric fields or query relevancy scores. Range filter over Function Query results Date Math - specify dates relative to "NOW" in queries and updates Dynamic search results clustering using Carrot2 Numeric field statistics such as min, max, average, standard deviation Combine queries derived from different syntaxes Auto-suggest functionality for completing user queries Allow configuration of top results for a query, overriding normal scoring and sorting Simple join capability between two document types Performance Optimizations Core Dynamically create and delete document collections without restarting Pluggable query handlers and extensible XML data format Pluggable user functions for Function Query Customizable component based request handler with distributed search support Document uniqueness enforcement based on unique key field Duplicate document detection, including fuzzy near duplicates Custom index processing chains, allowing document manipulation before indexing User configurable commands triggered on index changes Ability to control where docs with the sort field missing will be placed "Luke" request handler for corpus information Caching Configurable Query Result, Filter, and Document cache instances Pluggable Cache implementations, including a lock free, high concurrency implementation Cache warming in background When a new searcher is opened, configurable searches are run against it in order to warm it up to avoid slow first hits. During warming, the current searcher handles live requests. Autowarming in background The most recently accessed items in the caches of the current searcher are re-populated in the new searcher, enabling high cache hit rates across index/searcher changes. Fast/small filter implementation User level caching with autowarming support SolrCloud Centralized Apache ZooKeeper based configuration Automated distributed indexing/sharding - send documents to any node and it will be forwarded to correct shard Near Real-Time indexing with immediate push-based replication (also support for slower pull-based replication) Transaction log ensures no updates are lost even if the documents are not yet indexed to disk Automated query failover, index leader election and recovery in case of failure No single point of failure Admin Interface Comprehensive statistics on cache utilization, updates, and queries Interactive schema browser that includes index statistics Replication monitoring SolrCloud dashboard with graphical cluster node status Full logging control Text analysis debugger, showing result of every stage in an analyzer Web Query Interface w/ debugging output Parsed query output Lucene explain() document score detailing Explain score for documents outside of the requested range to debug why a given document wasn't ranked higher.