Introduction
Piano Contenttracks user interactions on all pages that include the Piano Insight JavaScript snippet. Piano then crawls, analyzes, and indexes this content. The search index serves recommendations.
A lot of these pages may contain content that should not come back as recommendations. The content to exclude is different for different use cases, but typical examples are front pages, section fronts, image galleries, expired events, etc. There is a variety of approaches available to blacklist, depending on the volume and extent of the exclusions.
Exclusion Strategies
Front pages
The crawler performs automatic detection of front and section pages. Pages identified as front pages will not be used for recommendations.
HTML Markup to explicitly exclude content
The recommended approach to exclude content from recommendations is to mark up the HTML of the pages with the cXenseParse:recs:recommendable tag as described in the Piano Content - Review and Refinement section.
If this markup is added to a page that has already been crawled, it will only take effect after the next crawl of that page. The /profile/content/push API can be used to push crawling of the url more quickly.
HTML Markup to define an expiration date
Piano Content recommendation automatically excludes expired content. To set the expiration date, mark up the HTML of the pages with the cXenseParserecs:expirationtime tag as described in the Piano Content - Review and Refinement section.
Query Filtering
If the sections of excluded content frequently change or are difficult to update, or if different Piano Content widgets will deliver different parts of the content, exclusions can also be done as part of the content configuration. This is done by appending an "and not" fragment to the content-configuration query parameter as described in the Content settings object.
|
|
Negative Sentiment Analysis
For several languages, we also have automatic Sentiment Analysis available. This will insert a sentiment key with the value "negative" for articles that are considered negative, and can be used to filter these from the result by adding the following to the query:
|
For deleted content
If the page returns a 404 or other error code in the HTTP response, the document will be deleted from the search index and will no longer be available for recommendations. As only URLs with traffic are automatically crawled, it could be necessary to push a crawl of the page using the /profile/content/push API. It is also possible to explicitly delete a page from the search index using the /profile/content/delete. However, if the page still exists and doesn't give an HTTP status error code (>400), it will be inserted in the index when it is next crawled.