Evaluating Website Classification API Accuracy for Crime & Justice
Your team needs to reliably detect and control traffic related to crime and justice topics—whether to block it for policy reasons, audit where ads run, or enrich a signup flow with risk signals. In this guide you’ll plug Klazify’s categorize endpoint into your pipeline, get a working yes/no decision for a “Crime & Justice” policy, and learn how to scale it with batching, caching, and fallback handling.
Why Klazify excels at classifying crime- and justice-related websites
When you screen domains that touch on policing, courts, investigations, or incident reporting, it’s not enough to inspect page titles or homepages. Klazify’s content analysis looks across the site and maps it to standardized categories, giving you consistent signals you can act on at ingestion time.
- Accurate Website Categorization Using AI: Klazify analyzes page content, structure, and signals rather than relying only on metadata. That helps it recognize nuanced content that overlaps news, government, and public-safety topics.
- Global Coverage: Crime and justice content appears in multiple languages. Klazify’s analysis supports multi-lingual sites so your enforcement or enrichment logic works globally.
- Real-Time Classification: Fresh classifications matter for rapidly evolving topics. Request-level analysis prevents stale inferences from derailing a decision.
- Industry-Level Categories: Results are mapped to IAB-style categories, which product and ad teams can align with existing brand-safety or compliance controls.
- Simple API Integration: A single REST call to the categorize endpoint returns categories, company data, logo, social signals, and related domains—useful for both filtering and enrichment.
- Compliance and Filtering: Precise categories simplify allow/block logic for law enforcement, legal, or sensitive crime-related pages, and can also power structured reviews rather than blanket blocks.
Concrete scenarios and a working path to deployment
Here are the common jobs your team can ship after integrating Klazify:
- Network or app blocking: Decide whether to allow or block access to domains that match your crime-and-justice criteria.
- Brand-safety auditing: Inspect where ads are placed and flag publishers or article sections that meet your risk definition.
- Signup and CRM enrichment: Add a “sensitive-topic” boolean and site metadata to profiles associated with a domain.
The rest of this guide shows you how to make a POST to the categorize endpoint, parse the results, map categories to your policy, and run it at scale with caching and batching.
Make a POST request to the categorize endpoint
Send the URL you want to evaluate in the request body. Use Bearer authentication with your API key.
curl -X POST "https://www.klazify.com/api/categorize" \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
--data '{"url":"https://www.justice.gov"}'
Replace YOUR_API_KEY with your actual key and set the url to any page or domain you want to categorize. You can send full URLs (page-level classification) or base domains (domain-level signals).
Below is a complete JSON response sample. Use it to wire up your parser and decision logic. Values will vary for your requests.
Understand the response and the fields you’ll use
Klazify returns categories plus additional domain intelligence you can plug into brand safety, enrichment, and analytics pipelines.
{
"domain": {
"categories": [
{
"confidence": 0.92,
"name": "/Computers & Electronics/Consumer Electronics",
"IAB-632-596": "Consumer Electronics/Technology & Computing/Consumer Electronics"
},
{
"confidence": 0.89,
"name": "/Internet & Telecom/Mobile & Wireless/Mobile Phones"
}
],
"social_media": null,
"logo_url": "https://klazify.s3.amazonaws.com/2110787991611585019600ed5fb1d1300.04730104.png"
},
"success": true,
"objects": {
"company": {
"url": "https://www.apple.com/",
"name": "Apple",
"city": "Cupertino",
"stateCode": "CA",
"countryCode": "US",
"employeesRange": "100K+",
"revenue": 274515000000,
"raised": null,
"tags": [
"E-commerce",
"Consumer Electronics",
"Mobile",
"B2C"
],
"tech": [
"omniture_adobe_analytics",
"atlassian_confluence",
"successfactors",
"apache_apex",
"talend",
"oracle_peoplesoft",
"salesforce",
"stripe",
"dell_boomi_atomsphere",
"gigya",
"sage_50cloud",
"quickbooks",
"webmethods",
"apache_tomcat",
"alteryx",
"tibco_rendezvous",
"atlassian_jira",
"..."
]
}
},
"domain_registration_data": {
"domain_age_date": "1987-02-19",
"domain_age_days_ago": "13026",
"domain_expiration_date": "2030-02-20",
"domain_expiration_days_left": "123"
},
"similar_domains": [
"bestbuy.com",
"icloud.com",
"microsoft.com",
"macrumors.com",
"google.com",
"samsung.com",
"twitter.com",
"hp.com",
"bhphotovideo.com",
"dell.com"
]
}
Key fields you’ll integrate:
- domain.categories: Array of category objects. Use name for readable labels, confidence for thresholding, and any IAB mapping key included for standardized taxonomy alignment.
- domain.logo_url: Useful for admin UI, CRM enrichment, and reviewer tooling.
- objects.company: Optional enriched company details (e.g., name, location, tags, tech) to augment risk scoring and customer segmentation.
- domain_registration_data: Age and expiration details can help contextualize new/unknown domains.
- similar_domains: Related domains for expansion, auditing, or sampling.
For crime-and-justice decisions, your filter will primarily operate on domain.categories[].name and any present IAB mapping fields. You can tighten or loosen decisions with confidence, apply thresholds, and escalate edge cases to review.
Example JSON responses you can test against
Use the same canonical sample to validate your parser, idempotence, and error handling across multiple stages (ingest, transform, decision). Keep the parsing logic identical in prod—only the input URL changes.
Parser validation: baseline
{
"domain": {
"categories": [
{
"confidence": 0.92,
"name": "/Computers & Electronics/Consumer Electronics",
"IAB-632-596": "Consumer Electronics/Technology & Computing/Consumer Electronics"
},
{
"confidence": 0.89,
"name": "/Internet & Telecom/Mobile & Wireless/Mobile Phones"
}
],
"social_media": null,
"logo_url": "https://klazify.s3.amazonaws.com/2110787991611585019600ed5fb1d1300.04730104.png"
},
"success": true,
"objects": {
"company": {
"url": "https://www.apple.com/",
"name": "Apple",
"city": "Cupertino",
"stateCode": "CA",
"countryCode": "US",
"employeesRange": "100K+",
"revenue": 274515000000,
"raised": null,
"tags": [
"E-commerce",
"Consumer Electronics",
"Mobile",
"B2C"
],
"tech": [
"omniture_adobe_analytics",
"atlassian_confluence",
"successfactors",
"apache_apex",
"talend",
"oracle_peoplesoft",
"salesforce",
"stripe",
"dell_boomi_atomsphere",
"gigya",
"sage_50cloud",
"quickbooks",
"webmethods",
"apache_tomcat",
"alteryx",
"tibco_rendezvous",
"atlassian_jira",
"..."
]
}
},
"domain_registration_data": {
"domain_age_date": "1987-02-19",
"domain_age_days_ago": "13026",
"domain_expiration_date": "2030-02-20",
"domain_expiration_days_left": "123"
},
"similar_domains": [
"bestbuy.com",
"icloud.com",
"microsoft.com",
"macrumors.com",
"google.com",
"samsung.com",
"twitter.com",
"hp.com",
"bhphotovideo.com",
"dell.com"
]
}
Decision engine unit test: multiple categories present
{
"domain": {
"categories": [
{
"confidence": 0.92,
"name": "/Computers & Electronics/Consumer Electronics",
"IAB-632-596": "Consumer Electronics/Technology & Computing/Consumer Electronics"
},
{
"confidence": 0.89,
"name": "/Internet & Telecom/Mobile & Wireless/Mobile Phones"
}
],
"social_media": null,
"logo_url": "https://klazify.s3.amazonaws.com/2110787991611585019600ed5fb1d1300.04730104.png"
},
"success": true,
"objects": {
"company": {
"url": "https://www.apple.com/",
"name": "Apple",
"city": "Cupertino",
"stateCode": "CA",
"countryCode": "US",
"employeesRange": "100K+",
"revenue": 274515000000,
"raised": null,
"tags": [
"E-commerce",
"Consumer Electronics",
"Mobile",
"B2C"
],
"tech": [
"omniture_adobe_analytics",
"atlassian_confluence",
"successfactors",
"apache_apex",
"talend",
"oracle_peoplesoft",
"salesforce",
"stripe",
"dell_boomi_atomsphere",
"gigya",
"sage_50cloud",
"quickbooks",
"webmethods",
"apache_tomcat",
"alteryx",
"tibco_rendezvous",
"atlassian_jira",
"..."
]
}
},
"domain_registration_data": {
"domain_age_date": "1987-02-19",
"domain_age_days_ago": "13026",
"domain_expiration_date": "2030-02-20",
"domain_expiration_days_left": "123"
},
"similar_domains": [
"bestbuy.com",
"icloud.com",
"microsoft.com",
"macrumors.com",
"google.com",
"samsung.com",
"twitter.com",
"hp.com",
"bhphotovideo.com",
"dell.com"
]
}
Enrichment pipeline test: using company, registration, and logo
{
"domain": {
"categories": [
{
"confidence": 0.92,
"name": "/Computers & Electronics/Consumer Electronics",
"IAB-632-596": "Consumer Electronics/Technology & Computing/Consumer Electronics"
},
{
"confidence": 0.89,
"name": "/Internet & Telecom/Mobile & Wireless/Mobile Phones"
}
],
"social_media": null,
"logo_url": "https://klazify.s3.amazonaws.com/2110787991611585019600ed5fb1d1300.04730104.png"
},
"success": true,
"objects": {
"company": {
"url": "https://www.apple.com/",
"name": "Apple",
"city": "Cupertino",
"stateCode": "CA",
"countryCode": "US",
"employeesRange": "100K+",
"revenue": 274515000000,
"raised": null,
"tags": [
"E-commerce",
"Consumer Electronics",
"Mobile",
"B2C"
],
"tech": [
"omniture_adobe_analytics",
"atlassian_confluence",
"successfactors",
"apache_apex",
"talend",
"oracle_peoplesoft",
"salesforce",
"stripe",
"dell_boomi_atomsphere",
"gigya",
"sage_50cloud",
"quickbooks",
"webmethods",
"apache_tomcat",
"alteryx",
"tibco_rendezvous",
"atlassian_jira",
"..."
]
}
},
"domain_registration_data": {
"domain_age_date": "1987-02-19",
"domain_age_days_ago": "13026",
"domain_expiration_date": "2030-02-20",
"domain_expiration_days_left": "123"
},
"similar_domains": [
"bestbuy.com",
"icloud.com",
"microsoft.com",
"macrumors.com",
"google.com",
"samsung.com",
"twitter.com",
"hp.com",
"bhphotovideo.com",
"dell.com"
]
}
Use these as fixture data for unit tests to confirm your parser extracts category names, confidences, and optional IAB-mapping keys without brittle assumptions.
Map categories to a yes/no decision for your “Crime & Justice” policy
Implement a deterministic mapping from the returned category names (and any IAB-mapped values present in the objects) to your internal policy. Many teams maintain a list of allowed and disallowed category path substrings and evaluate with confidence thresholds plus keyword backstops.
Python example: enforce an allow/block decision
import json
import requests
from typing import Dict, Any, List
API_URL = "https://www.klazify.com/api/categorize"
API_KEY = "YOUR_API_KEY"
# Configure your policy:
# - Use lowercase substring checks to avoid exact-path brittleness.
# - Keep this list versioned alongside policy reviews.
KEYWORDS_CRIME_JUSTICE = [
"crime", "criminal", "law enforcement", "police", "court", "justice", "prosecution",
"forensic", "prison", "incarceration", "victim", "investigation"
]
MIN_CONFIDENCE = 0.6 # treat categories at or above this as strong signals
def is_crime_justice(categories: List[Dict[str, Any]]) -> bool:
"""
Returns True if any category indicates crime/justice-related content.
Strategy:
1) If any category name includes policy keywords and has sufficient confidence.
2) If any taxonomy-mapped value (e.g., keys that look like 'IAB-...') includes keywords.
"""
for cat in categories:
name = (cat.get("name") or "").lower()
conf = float(cat.get("confidence") or 0.0)
# Check human-readable category path
if any(kw in name for kw in KEYWORDS_CRIME_JUSTICE) and conf >= MIN_CONFIDENCE:
return True
# Check any taxonomy mapping fields on the same object
for k, v in cat.items():
if k.startswith("IAB-"):
mapped = (v or "").lower()
if any(kw in mapped for kw in KEYWORDS_CRIME_JUSTICE) and conf >= MIN_CONFIDENCE:
return True
return False
def categorize(url: str) -> Dict[str, Any]:
headers = {"Authorization": f"Bearer {API_KEY}", "Content-Type": "application/json"}
payload = {"url": url}
resp = requests.post(API_URL, headers=headers, data=json.dumps(payload), timeout=20)
resp.raise_for_status()
return resp.json()
def decide(url: str) -> Dict[str, Any]:
data = categorize(url)
categories = data.get("domain", {}).get("categories", []) or []
decision = "block" if is_crime_justice(categories) else "allow"
return {
"url": url,
"decision": decision,
"categories": categories,
"logo_url": data.get("domain", {}).get("logo_url"),
"company": data.get("objects", {}).get("company"),
"domain_registration_data": data.get("domain_registration_data"),
"similar_domains": data.get("similar_domains")
}
if __name__ == "__main__":
result = decide("https://www.justice.gov")
print(json.dumps(result, indent=2))
This example shows how to:
- Call the categorize endpoint with Bearer authentication.
- Inspect domain.categories[].name and any IAB-mapped fields present per category object.
- Apply a confidence threshold and keyword-based mapping to an allow/block decision.
- Return logo, company, domain registration, and similar domains for downstream UI or review queues.
Operational details: accuracy signals and how to handle edge cases
Classifying content that touches on legal or public-safety topics often requires careful calibration. Here are practical techniques to preserve precision while maintaining coverage and latency.
- Confidence-aware thresholds: Use confidence on each category to qualify matches. Set a baseline threshold (for example, 0.6 as shown) and route low-confidence matches to review.
- Multi-signal AND/OR: Combine name substring checks with the presence of any IAB-mapped value on a category object. If both agree, you can be more certain.
- Negative keyword filters: If your policy excludes unrelated homonyms, maintain an exclusion list and require the absence of those terms before blocking.
- Sampling review: Persist the full JSON response when a decision changes for a given domain to support sampling and QA.
Caching, batching, and throughput
Most projects see repeated domains over time. Cache by normalized domain (e.g., example.com) and optionally by full path if you rely on page-level classification. Choose a TTL that suits your risk tolerance and how quickly the site’s content changes.
- Caching keys: Use a consistent canonical form. For URLs, consider both domain-level and path-level keys if your policy depends on specific sections.
- Cache invalidation: Refresh on 4xx/5xx retries, on policy changes, or after significant content changes uncovered by your crawler.
- Batching: Group requests within your job scheduler to minimize overhead and respect rate limits. If you maintain a queue, de-duplicate domains before enqueueing.
- Timeouts and retries: Use bounded timeouts, exponential backoff, and idempotent job design.
For rate limits and quotas, consult your plan details and ensure backpressure in your pipeline. If you expect bursts (e.g., ad-slot scanning), add a short-lived cache layer in front of the API to coalesce duplicate requests.
Handling unknown or new domains
When a domain appears new or lacks strong categorical signals, apply a measured fallback:
- Use domain_registration_data to weigh new or about-to-expire domains differently in your policy.
- Escalate “unknown” or low-confidence categories to manual review or a secondary crawl pass.
- Leverage similar_domains to expand your review set and infer context while avoiding automatic propagation of a risky decision.
Mapping Klazify categories to your internal taxonomy
Many organizations maintain their own policy taxonomy. The domain.categories[].name field provides a readable path, while any present IAB-style mapping within the category object lets you align with standardized labels. A simple mapping layer can:
- Normalize Klazify category names into your internal enums.
- Map one or more Klazify paths into a single internal “Crime & Justice” bucket.
- Version mappings alongside release cycles so audits can reproduce historical decisions.
Recommended pattern:
- Create a mapping table from category substrings to your internal policy buckets.
- Use confidence thresholds and optional IAB-mapped fields for tie-breaking.
- Emit both the raw Klazify categories and your normalized label to your data warehouse for QA.
Putting it all together in your pipeline
Here’s how teams typically deploy this in production:
- Normalization: Canonicalize incoming URLs/domains and check the cache.
- Lookup: If uncached, call POST /api/categorize with the URL in the body.
- Parse: Extract domain.categories[], logo_url, objects.company, and other fields your use case needs.
- Decide: Run the category-to-policy mapper to set allow/block or risk flags.
- Persist: Store the JSON payload, derived labels, and decision for debugging and audit.
- Monitor: Track decision distribution and confidence histograms; sample low-confidence edge cases for review.
Field-by-field tips for crime-and-justice use cases
- domain.categories[].name: Treat as the primary decision feature. It’s hierarchical; substring checks work well for grouping related paths.
- domain.categories[].confidence: Use to set routing rules (auto-allow/auto-block/review).
- domain.categories[].IAB-...: If included, use as a standardized signal to reduce overfitting to a single path naming convention.
- objects.company.tags: Adds context for B2B flows (e.g., whether a legal-services firm appears). Do not rely solely on tags to determine sensitive-content policy.
- similar_domains: Expand your sample during testing to see how decisions generalize within a publisher network or organizational family.
Governance, auditing, and versioning
Policy changes can have wide impact. Version your mapping table, record the Klazify response payload alongside your derived label, and store the code version that produced the decision. This enables reproducibility and targeted rollbacks if a specific pattern yields false positives.
- Schema stability: Keep the raw JSON in cold storage to re-derive labels later.
- Determinism: Implement pure functions for mapping category arrays to your internal policy.
- Explainability: When you block, include which category name and confidence triggered the decision.
Where to go from here
Integrate the single endpoint shown above, align the response fields with your internal policy, and then iterate on thresholds with sampled reviews. See the complete reference and create a free account to start testing.
FAQ
How do I reduce false positives for borderline pages?
Raise the confidence threshold, require both a category name match and a taxonomy-mapped match when available, and add a short list of negative keywords. Route the remainder to a manual review queue.
Should I classify at the domain or URL level?
If sections within a site vary significantly, send page URLs for finer-grained control. If your policy is domain-wide, cache by domain to reduce calls.
What should I cache, and for how long?
Cache the full JSON response keyed by the normalized URL or domain, plus your derived label. Set TTLs based on your risk tolerance and how often site content changes. Refresh on policy updates or when you detect content shifts.
How do I handle unknown or very new domains?
Use domain_registration_data as a signal, lower the confidence threshold for escalation (not auto-block), and capture similar_domains for contextual review.
Can I align Klazify categories with my existing taxonomy?
Yes. Map domain.categories[].name and any IAB-mapped values to your internal enums via a lookup table, then store both raw and normalized labels to support auditing.
Create a free account on Try Klazify API for free, call the categorize endpoint with your first set of domains, and wire the returned categories, confidence, and optional taxonomy mappings into your crime-and-justice policy for blocking, brand safety, or enrichment.
Ready to use Klazify?
Start classifying websites, enriching company data, and exploring web intelligence.
Get Started Free