Metric Formulas and Calibration

Every TechEco DZ metric (0-10 or 0-100), with its exact formula: review signals, digital maturity, popularity, economic geography, tech-hiring potential and data confidence.

Review signals

company.review_volume · landmark.review_volume · company.rating_quality · landmark.rating_quality

Review volume. The raw Google review count is transformed, then robustly normalized within the city, where p05/p95 are the 5th/95th percentiles observed among the city's rated entities. If p95 equals p05 (no spread), every rated entity scores 5.

Rating quality. The raw star rating (0 to 5) is first shrunk toward the city mean by sample strength: a handful of 5-star reviews can't outrank thousands of 4.6-star ones.

The result is then scaled to 0-10 through a logistic function centered on the mean:

Missing reviews. An entity with zero reviews gets the neutral average of both scores above across the city's rated entities, flagged imputed = true with reduced confidence (0.35 instead of 1.0). An imputed estimate is never presented as an actual Google rating.

Digital maturity

company.digital_maturity

An additive 0-10 score, counted only when the declared URL is a genuinely owned domain (not a social-media page) and was successfully crawled:

  • +2.5 reachable, verified owned domain
  • +1.0 successful crawl
  • +1.0 an "about" or "services" page was found
  • +1.5 a "careers" page was found
  • +0.5 a "contact" page was found
  • +1.0 served over HTTPS
  • +1.0 crawled within the last 90 days

The social-profile bonus applies independent of the website and caps out at 3 validated profiles.

Popularity

company.public_prominence · landmark.intrinsic

The score shown everywhere as "Popularity" is a weighted average of the metrics above, specific to the entity type:

Because every component is still computed without reviews (see "Review signals" above), popularity is always computed for every published entity. A low score means genuinely low popularity, not missing data.

Economic geography

landmark.economic_relevance · landmark.ecosystem_relevance · company.economic_location

Economic relevance (landmarks). A fixed weight by landmark type, not a computed metric: factories and industrial zones = 10; shopping malls and business districts = 9; transit hubs, commercial hubs and markets = 8; universities and hospitals = 7; hotels, government buildings, busy roads/squares, ports, upscale residential = 6; beach resorts and airports = 5; any other type = 4.

Ecosystem relevance (landmarks). The count of companies within 3km of the landmark, robustly normalized (same p05-p95 method as review volume) across every landmark in the city. A relative measure of local economic density.

Economic location (companies). For each company, the distance-weighted sum of every landmark's popularity in the city:

A landmark 3km away contributes about 37% of its popularity, one 6km away about 13%. The result is then robustly normalized (p05-p95) across the city's companies.

Tech-hiring potential

company.tech_opportunity · company.top_role_potential · tech_calibration (0-100)

This potential is produced in two distinct, versioned stages, never from a single language-model guess.

Stage 1, assessment. A structured LangChain agent assesses all 16 controlled specialties and declares, per specialty, a model score (0-100) and a self-declared evidence level (none, weak, moderate, strong or direct). That self-declared level is then capped by a deterministic check (specialty terms searched for in the company's name, Maps category, sector/subsector and observed job postings): with no match, the level is capped to "weak" (or "moderate" for a SoftwareIT-sector company). Each specialty's final score:

An unsupported score can therefore only move partway from the sector baseline.

Stage 2, calibration. Every specialty is then re-checked against evidence actually stored for the company (job postings, crawled pages, social profiles, Maps category, community reviews — see below), never against the model's own claims. The best matching source sets a reliability and a score cap:

A specialty with no grounding at all can therefore never exceed 49/100.

Community-review signal. A signed-in user leaving a review can tag the specialties they believe the company is hiring for. That tag alone proves nothing — it only becomes evidence once 2 separate reviews name the same specialty (a "weak" cap, 59/100), or 5 separate reviews (a "moderate" cap, 74/100) — never beyond that, no matter how many reviews pile up. When a review includes a comment, a language model first verifies the comment text actually corroborates the tagged specialty (rejecting vague or contradicting comments, e.g. mentioning a layoff); a tag with no comment is counted as-is. Real, scraped evidence (a job posting, a careers page) always outranks this community signal.

Company-level score (shown as a percentage, never presented as a vacancy), where top1-3 are the three highest calibrated scores:

company.tech_opportunity (out of 10) is this same opportunity divided by 10; company.top_role_potential is the highest-ranked specialty's calibrated score.

Hiring activity

company.hiring_activity · company.career_infrastructure

Hiring activity. A single stale posting decays toward 0; several recent postings push the score toward 10. Capped at 10.

Career infrastructure. +6 if a distinct "careers" page was found on the site, +4 more if at least one active posting was observed. The maximum (10) requires both.

Data confidence

company.data_confidence · company.entity_quality_risk

Data confidence, an internal quality signal, not a popularity score:

where reviews_present is 1 if reviews are not imputed (else 0.35), and tech_evidence_weight reuses the reliability scale above for the company's best tech-evidence level.

Entity quality risk, an internal QA-only signal, never shown publicly. If two or more published companies share the exact same rounded coordinates (7 decimal places, sub-meter precision, a strong signal of a duplicate or miscoded location):

Otherwise the risk is 0.