URL Classification API Rules for Crime & Justice Categorization

October 03, 2026
URL Classification API Rules for Crime & Justice Categorization

You need to automatically identify “Crime & Justice” content across the web so you can block it in a network filter, enforce brand-safety rules for ad placements, or enrich signups in your CRM with a risk flag. By the end of this guide, you’ll send a URL to Klazify’s categorize endpoint, read the result, and turn it into a yes/no decision for your “Crime & Justice” policy—then scale it with caching, batching, and auditing.

Why Klazify excels at “Crime & Justice” URL classification

Content tagged as crime, justice, or legal systems often spans government portals, policy pages, news reporting, court records, NGO resources, and advocacy sites. These pages change frequently and appear in many languages. Klazify’s approach aligns well with this domain for several reasons:

Illustration: URL Classification API Rules for Crime & Justice Categorization
  • Accurate website categorization using AI: Klazify analyzes full page content, not just metadata. That helps separate legal resources (e.g., court information) from unrelated topics that simply mention justice-related terms in passing.
  • Global coverage: Crime, legal, and justice content exists in multilingual contexts (e.g., national justice departments, courts, NGOs). Klazify’s global coverage helps classify such content consistently.
  • Real-time classification: Justice-related content evolves with policy updates, court decisions, and investigations. Klazify supports fresh analysis to reflect those changes rather than relying on stale snapshots.
  • Industry-level categories (IAB taxonomy): Results map to an industry taxonomy, simplifying alignment with ad tech, compliance, and filtering pipelines.
  • Simple API integration: The single REST endpoint is straightforward to call from your backend or data pipelines.
  • Compliance and filtering: The response includes structured categories and confidence values you can apply to allow/deny lists, escalation workflows, and audit logs.

If your team needs to flag or filter justice-related content in real time, start by creating a free account at Register and review the Documentation.

The scenario: build a “Crime & Justice” policy you can enforce

Imagine you run three workflows:

  • Brand safety: Before running a campaign, you verify that publisher pages don’t fall under your “Crime & Justice” exclusion set.
  • Network filtering: On a corporate or public Wi-Fi, you block or restrict access to sites that match your “Crime & Justice” category rules.
  • Lead enrichment: When a company signs up, you classify its website to flag accounts that operate in legal systems you want to route to a specialized team.

All three rely on the same decision: map Klazify’s category output to your own “Crime & Justice” rule set, then return a boolean allow/block or a risk score. The rest of this article shows how.

How the categorize endpoint works

Send a URL to the categorize endpoint, receive a JSON payload with category paths, confidence, and additional domain and company fields. You’ll use:

  • domain.categories[].name: The hierarchical category path you’ll compare with your policy list.
  • domain.categories[].confidence: A float (0 to 1) you can threshold for stricter decisions.
  • Optional enrichment: domain.logo_url, objects.company fields, domain_registration_data, and similar_domains for downstream analytics, whitelists, or review queues.

Reference the MCP if you need a high-level capability profile before integrating.

Make a request with curl (POST /api/categorize)

Below is a complete, copy-pasteable example using Bearer auth and a URL in the JSON body. Replace YOUR_API_KEY with your token.

curl -sS -X POST "https://www.klazify.com/api/categorize" \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"url":"https://www.justice.gov/"}'

Tip: For page-level checks on news or legal articles, send the full URL. For domain-level allow/deny, send the homepage URL. Cache per registrable domain to avoid repeat calls (details below).

The official JSON response format (unchanged)

This is the canonical sample response. Use it to understand structure and field names. Do not alter field names or values when integrating.


{
"domain": {
"categories": [
{
"confidence": 0.92,
"name": "/Computers & Electronics/Consumer Electronics",
"IAB-632-596": "Consumer Electronics/Technology & Computing/Consumer Electronics"
},
{
"confidence": 0.89,
"name": "/Internet & Telecom/Mobile & Wireless/Mobile Phones"
}
],
"social_media": null,
"logo_url": "https://klazify.s3.amazonaws.com/2110787991611585019600ed5fb1d1300.04730104.png"
},
"success": true,
"objects": {
"company": {
"url": "https://www.apple.com/",
"name": "Apple",
"city": "Cupertino",
"stateCode": "CA",
"countryCode": "US",
"employeesRange": "100K+",
"revenue": 274515000000,
"raised": null,
"tags": [
"E-commerce",
"Consumer Electronics",
"Mobile",
"B2C"
],
"tech": [
"omniture_adobe_analytics",
"atlassian_confluence",
"successfactors",
"apache_apex",
"talend",
"oracle_peoplesoft",
"salesforce",
"stripe",
"dell_boomi_atomsphere",
"gigya",
"sage_50cloud",
"quickbooks",
"webmethods",
"apache_tomcat",
"alteryx",
"tibco_rendezvous",
"atlassian_jira",
"..."
]
}
},
"domain_registration_data": {
"domain_age_date": "1987-02-19",
"domain_age_days_ago": "13026",
"domain_expiration_date": "2030-02-20",
"domain_expiration_days_left": "123"
},
"similar_domains": [
"bestbuy.com",
"icloud.com",
"microsoft.com",
"macrumors.com",
"google.com",
"samsung.com",
"twitter.com",
"hp.com",
"bhphotovideo.com",
"dell.com"
]
}

What fields you’ll actually use for “Crime & Justice” policy

Even though the example above shows technology-related categories, the same fields appear when a page matches your “Crime & Justice”-relevant topics. Here’s how to interpret them:

  • domain.categories[].name: Hierarchical category string you compare to your curated “Crime & Justice” list (maintained by your team).
  • domain.categories[].confidence: If borderline, you can set a minimum threshold (e.g., only block when ≥ a chosen value). This field is unitless; treat it as a probability-like score from 0.0 to 1.0.
  • domain.categories[].IAB-…: When present, this maps to a structured IAB taxonomy path. Use it if your policy is IAB-centric rather than free-text category paths.
  • objects.company.*: Use for enrichment (e.g., CRM fields, routing rules for accounts that operate within legal ecosystems).
  • domain_registration_data.*: Helpful for fraud/risk signals or QA (e.g., very young domains might merit review).
  • similar_domains[]: Useful for precomputing allow/deny lists or bulk audits of related properties.

Map categories to a yes/no decision in Python

This minimal function calls Klazify, extracts categories, and returns True if any category matches your organization’s “Crime & Justice” list. Maintain that list centrally (e.g., via a config service) so policy changes don’t require redeploys.

import os
import json
import time
import requests
from urllib.parse import urlparse

KLAZIFY_API_KEY = os.getenv("KLAZIFY_API_KEY", "YOUR_API_KEY")
KLAZIFY_URL = "https://www.klazify.com/api/categorize"

# Your curated policy lists. Keep these in configuration, not code.
CRIME_JUSTICE_CATEGORY_PATHS = set([
# Insert your approved category paths here that indicate "Crime & Justice".
# Example: "/Law & Government/..." (maintain your own exact strings)
])

IAB_CODES_FOR_CRIME_JUSTICE = set([
# If you use IAB mapping, place the exact IAB code strings here.
])

MIN_CONFIDENCE = 0.80 # adjust based on QA; ranges from 0.0 to 1.0

def registrable_domain(url: str) -> str:
"""
Extract a cache key like example.com from a URL.
Keep your own PSL/ETLD+1 logic if needed.
"""
hostname = urlparse(url).hostname or ""
parts = hostname.split(".")
if len(parts) >= 2:
return ".".join(parts[-2:])
return hostname

def categorize(url: str) -> dict:
headers = {
"Authorization": f"Bearer {KLAZIFY_API_KEY}",
"Content-Type": "application/json"
}
payload = {"url": url}
resp = requests.post(KLAZIFY_URL, headers=headers, data=json.dumps(payload), timeout=20)
resp.raise_for_status()
return resp.json()

def is_crime_justice(url: str, cache: dict, ttl_seconds: int = 86400) -> bool:
key = registrable_domain(url)
now = int(time.time())

# Return cached decision if fresh
if key in cache:
item = cache[key]
if now - item["ts"] < ttl_seconds:
return item["decision"]

data = categorize(url)

decision = False
domain_data = data.get("domain", {}) or {}
categories = domain_data.get("categories", []) or []

for cat in categories:
name = cat.get("name", "")
conf = float(cat.get("confidence", 0.0))

if conf < MIN_CONFIDENCE:
continue

# Match on category path
if name in CRIME_JUSTICE_CATEGORY_PATHS:
decision = True
break

# Optionally match on IAB code/path if present
for k, v in cat.items():
if k.startswith("IAB-"):
if v in IAB_CODES_FOR_CRIME_JUSTICE:
decision = True
break
if decision:
break

# Cache decision
cache[key] = {"decision": decision, "ts": now}
return decision

if __name__ == "__main__":
cache_store = {}
test_url = "https://www.justice.gov/"
print(is_crime_justice(test_url, cache_store))

Notes:

  • Cache by registrable domain (e.g., example.com) for homepage checks; consider URL-level caching for news sites with mixed content.
  • Use confidence to tighten or loosen blocking behavior. For sensitive filters, require a higher threshold.
  • Store your policy lists outside code (database, config file), and add change control for updates.

A second official JSON example for parsing practice (unchanged)

Use the same response to test your parsing logic for categories, IAB mapping, and enrichment fields.


{
"domain": {
"categories": [
{
"confidence": 0.92,
"name": "/Computers & Electronics/Consumer Electronics",
"IAB-632-596": "Consumer Electronics/Technology & Computing/Consumer Electronics"
},
{
"confidence": 0.89,
"name": "/Internet & Telecom/Mobile & Wireless/Mobile Phones"
}
],
"social_media": null,
"logo_url": "https://klazify.s3.amazonaws.com/2110787991611585019600ed5fb1d1300.04730104.png"
},
"success": true,
"objects": {
"company": {
"url": "https://www.apple.com/",
"name": "Apple",
"city": "Cupertino",
"stateCode": "CA",
"countryCode": "US",
"employeesRange": "100K+",
"revenue": 274515000000,
"raised": null,
"tags": [
"E-commerce",
"Consumer Electronics",
"Mobile",
"B2C"
],
"tech": [
"omniture_adobe_analytics",
"atlassian_confluence",
"successfactors",
"apache_apex",
"talend",
"oracle_peoplesoft",
"salesforce",
"stripe",
"dell_boomi_atomsphere",
"gigya",
"sage_50cloud",
"quickbooks",
"webmethods",
"apache_tomcat",
"alteryx",
"tibco_rendezvous",
"atlassian_jira",
"..."
]
}
},
"domain_registration_data": {
"domain_age_date": "1987-02-19",
"domain_age_days_ago": "13026",
"domain_expiration_date": "2030-02-20",
"domain_expiration_days_left": "123"
},
"similar_domains": [
"bestbuy.com",
"icloud.com",
"microsoft.com",
"macrumors.com",
"google.com",
"samsung.com",
"twitter.com",
"hp.com",
"bhphotovideo.com",
"dell.com"
]
}

How to use the returned fields for filtering, enrichment, and audits

Field Type Usage in “Crime & Justice” policy Example downstream action
domain.categories[].name String Compare against your curated list to flag Crime & Justice content. Block access; exclude from ad placements; tag CRM record.
domain.categories[].confidence Float (0.0–1.0) Threshold to avoid false positives; prompt review if borderline. Escalate to manual QA when within a gray band.
domain.categories[].IAB-… String Use IAB-mapped path if your taxonomy is IAB-based. Automate mapping to existing brand-safety exclusion lists.
objects.company.name, url String Enrich CRM with company identity operating the site. Assign to a specialist legal vertical team.
domain_registration_data.* String/Number Risk checks (e.g., very new domains) or data provenance. Queue suspicious domains for additional verification.
similar_domains[] Array Seed bulk audits or pre-warm caches for related properties. Scan similar domains before activating a campaign.

A third official JSON example for QA pipelines (unchanged)

Teams often keep a known-good sample alongside parsers and unit tests. Reuse this block in tests to ensure your code doesn’t break when you add fields to logs or BI flows.


{
"domain": {
"categories": [
{
"confidence": 0.92,
"name": "/Computers & Electronics/Consumer Electronics",
"IAB-632-596": "Consumer Electronics/Technology & Computing/Consumer Electronics"
},
{
"confidence": 0.89,
"name": "/Internet & Telecom/Mobile & Wireless/Mobile Phones"
}
],
"social_media": null,
"logo_url": "https://klazify.s3.amazonaws.com/2110787991611585019600ed5fb1d1300.04730104.png"
},
"success": true,
"objects": {
"company": {
"url": "https://www.apple.com/",
"name": "Apple",
"city": "Cupertino",
"stateCode": "CA",
"countryCode": "US",
"employeesRange": "100K+",
"revenue": 274515000000,
"raised": null,
"tags": [
"E-commerce",
"Consumer Electronics",
"Mobile",
"B2C"
],
"tech": [
"omniture_adobe_analytics",
"atlassian_confluence",
"successfactors",
"apache_apex",
"talend",
"oracle_peoplesoft",
"salesforce",
"stripe",
"dell_boomi_atomsphere",
"gigya",
"sage_50cloud",
"quickbooks",
"webmethods",
"apache_tomcat",
"alteryx",
"tibco_rendezvous",
"atlassian_jira",
"..."
]
}
},
"domain_registration_data": {
"domain_age_date": "1987-02-19",
"domain_age_days_ago": "13026",
"domain_expiration_date": "2030-02-20",
"domain_expiration_days_left": "123"
},
"similar_domains": [
"bestbuy.com",
"icloud.com",
"microsoft.com",
"macrumors.com",
"google.com",
"samsung.com",
"twitter.com",
"hp.com",
"bhphotovideo.com",
"dell.com"
]
}

Batching, rate limits, retries, and timeouts

When you need to classify at scale (e.g., nightly crawls or ad inventory scans), plan for:

  • Batching: Organize URLs by domain to leverage cache hits. If multiple pages share a domain and you only need domain-level policy, deduplicate before calling the API.
  • Concurrency: Use a worker pool with bounded concurrency to avoid exceeding your plan’s throughput. If you see HTTP 429 or related throttling, back off and retry with jitter.
  • Retries: Retry idempotently on transient network failures or 5xx responses with exponential backoff. Avoid retry storms by capping attempts.
  • Timeouts: Client-side timeouts protect your workers; 10–30 seconds is typical. Log slow responses separately for investigation.
  • Queueing: For very large lists, push URLs into a message queue and process asynchronously to smooth out spikes.

Always monitor overall success rates and average latency in your observability stack. If you need implementation details, see the Documentation.

Caching per domain and controlling TTLs

Caching reduces cost and latency:

  • Key by registrable domain for homepage checks; key by full URL for sites with diverse content per page (e.g., news articles).
  • Set a TTL (e.g., 24 hours) for most sites. Shorten TTLs for fast-changing properties, and lengthen for stable institutional domains.
  • Store the raw JSON plus your computed decision and relevant metadata (timestamp, confidence) for auditing.
  • Invalidate cache when your policy list changes to re-evaluate decisions.

Handling unknown or new domains

For domains you’ve never seen or just registered recently:

  • Classify on first encounter and store the raw response.
  • If domain_registration_data suggests a very new domain, you may choose stricter thresholds or queue for manual review.
  • If no categories are returned, set a fallback policy: allow, block, or “review required.” Document this default in your runbook.

Mapping Klazify output to your taxonomy

Some teams maintain a bespoke taxonomy. You can:

  • Create a mapping from domain.categories[].name to your internal tags. Keep the mapping in a table your policy engine can hot-reload.
  • Alternatively, standardize on IAB mapping provided in fields like IAB-… when present. Maintain an IAB-to-internal mapping table.
  • Version your mappings so you can reprocess historical records consistently.

Pipeline wiring: brand safety, filtering, enrichment

Once a URL is classified, you can execute different actions:

  • Brand safety: If a page matches your “Crime & Justice” exclusion, do not bid or place creatives. Log domain, categories, and the confidence used to make the decision.
  • Network filtering: Enforce block or warn pages. Emit an audit event that includes domain.categories[].name and confidence for compliance.
  • CRM enrichment: Attach a boolean field like is_crime_justice and optionally store raw response JSON in a secure data store for analysts.

JavaScript snippet (optional) to classify a single URL

If you prefer a Node.js utility for ad-hoc checks, this example mirrors the Python logic at a high level. Insert your own configuration for the policy lists.

import fetch from "node-fetch";

const API_KEY = process.env.KLAZIFY_API_KEY || "YOUR_API_KEY";
const ENDPOINT = "https://www.klazify.com/api/categorize";

const CRIME_JUSTICE_PATHS = new Set([
// Add your exact category paths here.
]);

const IAB_CRIME_JUSTICE = new Set([
// Add your exact IAB strings here if you match on IAB mapping.
]);

const MIN_CONFIDENCE = 0.80;

async function isCrimeJustice(url) {
const resp = await fetch(ENDPOINT, {
method: "POST",
headers: {
"Authorization": `Bearer ${API_KEY}`,
"Content-Type": "application/json"
},
body: JSON.stringify({ url })
});

if (!resp.ok) {
throw new Error(`HTTP ${resp.status}`);
}

const data = await resp.json();
const categories = (data.domain && data.domain.categories) || [];
for (const c of categories) {
const name = c.name || "";
const conf = typeof c.confidence === "number" ? c.confidence : 0.0;
if (conf < MIN_CONFIDENCE) continue;

if (CRIME_JUSTICE_PATHS.has(name)) return true;

for (const [k, v] of Object.entries(c)) {
if (k.startsWith("IAB-") && IAB_CRIME_JUSTICE.has(v)) {
return true;
}
}
}
return false;
}

isCrimeJustice("https://www.justice.gov/").then(console.log).catch(console.error);

Auditing ad inventory and publisher lists

When scanning many publishers before a campaign, parallelize requests with caching and maintain an audit table including:

  • input_url, resolved_domain
  • top category names (and IAB paths if present)
  • top confidence
  • decision flag and decision reason (e.g., “matched path in policy set”)
  • timestamp and cache status (HIT/MISS)

This structure lets you explain exclusions to partners and iterate on policy lists. For repeat campaigns on the same list, re-use cached results and only refresh domains older than your TTL.

Security and compliance notes

  • Keep API keys encrypted in transit and at rest. Rotate keys periodically.
  • Log only what you need for audits. Consider hashing user identifiers if you store traffic logs.
  • Review your default behavior for unknown categories and document escalation steps in your incident runbooks.

FAQ

Q: How do I handle pages where multiple categories are returned?
A: Evaluate each category against your policy. If any category matches your “Crime & Justice” list and meets your confidence threshold, treat the page as in-scope. Optionally, prefer the highest-confidence category.

Q: Should I classify by domain or by full URL?
A: Use domain-level classification for static, single-purpose sites. Use URL-level checks for dynamic publishers (e.g., news) where different pages can fall inside or outside your policy.

Q: How often should I refresh cached results?
A: Set a TTL based on volatility. Many teams use daily refreshes for general sites and shorter TTLs for fast-changing publishers. Invalidate caches when you update your policy lists.

Q: How can I map Klazify’s output to my internal taxonomy?
A: Maintain two mapping tables: one from category paths to your tags and one from IAB strings (when present) to your tags. Version the mappings and store them in a central config service.

Q: What if a domain is too new or returns minimal content?
A: Apply a fallback policy (allow/block/review) and queue for recheck later. Use domain_registration_data to inform stricter reviews for very young domains.

Start classifying URLs with Klazify and turn category outputs into enforceable rules across brand safety, filtering, and enrichment. Create your free account here: Register. For implementation details, see the Documentation and explore capabilities via the MCP.

Ready to use Klazify?

Start classifying websites, enriching company data, and exploring web intelligence.

Get Started Free