Skip to content

Instantly share code, notes, and snippets.

@eggplants
Forked from advenk/SPARQL.md
Last active September 24, 2026 17:30
Show Gist options
  • Select an option

  • Save eggplants/4b98b926577e46a8d8a739435773ecfa to your computer and use it in GitHub Desktop.

Select an option

Save eggplants/4b98b926577e46a8d8a739435773ecfa to your computer and use it in GitHub Desktop.
Hindi SPARQL Endpoint - Motivation and Sample Queries

Report: Hindi DBpedia SPARQL Endpoint Query Demonstration

Date: 29 July 2025

Objective: To demonstrate the successful deployment and operational capability of the new Hindi DBpedia SPARQL endpoint.

1. Introduction

This document first introduces how to deploy the Hindi endpoint on any server with simple docker commands. This endpoint is deployed against the Hindi wiki dump as of 1st June 2025. If this needs to be updated, the dump needs to be extracted again and deployed seperately.

This document then serves as a proof-of-concept to showcase the motivation and functionality of the newly deployed Hindi DBpedia knowledge graph. The following sections contain a series of SPARQL queries designed to validate the endpoint's ability to handle a wide range of tasks, from simple entity lookups to complex relationship traversals and data aggregations. All queries are executed against the temporary endpoint deployed at http://hi.dbpedia.org/sparql endpoint.

Deploying on Server

To deploy on server:

docker run -d -p 8890:8890 -p 1111:1111 --name hindi-sparql 42bitstogo/hindi-dbpedia-sparql:latest

2. Motivation and Benefits

The deployment of a dedicated Hindi DBpedia chapter is a significant initiative driven by the need to make vast, community-curated knowledge accessible in a structured, machine-readable format. The primary motivations and benefits include:

  • Unlocking Hindi Wikipedia: The Hindi Wikipedia is a rich and expansive source of information. However, as a collection of articles, its data is primarily unstructured prose intended for human readers. DBpedia transforms this repository into a structured knowledge graph, converting information from infoboxes, tables, and categories into queryable data.

  • Providing a Centralized Query Endpoint: Instead of requiring developers to build complex and fragile web scrapers to parse individual Wikipedia pages, the Hindi DBpedia provides a single, stable, and powerful SPARQL endpoint. This centralized graph acts as a unified source of truth for structured data from Hindi Wikipedia, dramatically simplifying data access for applications.

  • Empowering Localized Applications: By providing structured data in Hindi, this project empowers developers to build applications with deep cultural and regional relevance for one of the world's largest language communities. Potential applications include intelligent chatbots, localized search engine knowledge panels, academic research tools, and recommendation systems tailored to a Hindi-speaking audience.

  • Enhancing the Semantic Web: The Hindi DBpedia chapter contributes to the global Linked Open Data cloud. By using standardized ontologies and URIs, its data can be interlinked with other DBpedia language chapters and other datasets, fostering a more interconnected and multilingual web of data.


3. Sample queries

This section focuses on a single, well-defined entity to demonstrate fundamental query patterns.

3.1. Locating a Specific Entity by Name

  • Objective: To find the URI for the resource representing "Amitabh Bachchan" by searching for its Hindi name.
  • Methodology: This query searches for all resources of type dbo:Person and filters them for those whose Hindi language (hi) name contains the string "अमिताभ बच्चन".
PREFIX foaf: <http://xmlns.com/foaf/0.1/>
PREFIX dbo: <http://dbpedia.org/ontology/>

SELECT ?person ?name
WHERE {
  ?person a dbo:Person .
  ?person foaf:name ?name .
  FILTER(LANG(?name) = "hi")
  # Add this filter to search INSIDE the name string
  FILTER(CONTAINS(?name, "अमिताभ बच्चन"))
}
LIMIT 10

3.2. Retrieving All Properties of an Entity

  • Objective: To retrieve all predicate-object pairs for a known entity URI to inspect its available data.
  • Methodology: Given the specific URI for Amitabh Bachchan, this query selects all outgoing properties (?p) and their corresponding values (?o).
SELECT ?p ?o
WHERE {
  <http://hi.dbpedia.org/resource/अमिताभ_बच्चन> ?p ?o .
}
LIMIT 200

3.3. Traversing a Relationship to Find Linked Data

  • Objective: To follow a single relationship (dbo:parent) from a starting entity to find the URI of a related entity.
  • Methodology: The query starts at the Amitabh Bachchan resource and selects the object connected by the dbo:parent predicate.
PREFIX dbo: <http://dbpedia.org/ontology/>
PREFIX dbr-hi: <http://hi.dbpedia.org/resource/>

# Just find the parent's URI, nothing else.
SELECT ?parentURI
WHERE {
  dbr-hi:अमिताभ_बच्चन dbo:parent ?parentURI .
}

3.4. Inspecting a Related Entity's Data

  • Objective: To use the URI found in the previous step to retrieve all data for Harivansh Rai Bachchan.
  • Methodology: This demonstrates a common workflow where the output of one query (the parent's URI) becomes the input for a subsequent query.
# Let's see all the data for Harivansh Rai Bachchan
SELECT ?p ?o
WHERE {
  <http://hi.dbpedia.org/resource/हरिवंश_राय_बच्चन> ?p ?o .
}
LIMIT 200

3.5. Multi-Step Traversal

  • Objective: To find the death date of Amitabh Bachchan's parent by chaining two relationships together in a single query.
  • Methodology: This query first finds the parent's URI via the dbo:parent property and then, using that variable, finds the associated dbo:deathDate.
PREFIX dbo: <http://dbpedia.org/ontology/>
PREFIX dbr-hi: <http://hi.dbpedia.org/resource/>
PREFIX rdfs: <http://www.w3.org/2000/01/rdf-schema#>

SELECT ?parentDeathDate
WHERE {
  # 1. Start with Amitabh and find his parent's URI
  dbr-hi:अमिताभ_बच्चन dbo:parent ?parentURI .
  # 2. Use the parent's URI to find their birth date
  ?parentURI dbo:deathDate ?parentDeathDate .
}

4. Categorical and Aggregate Queries

This section demonstrates queries that operate on entire classes of data to perform counts, checks, and broad searches.

4.1. Counting Instances of a Type

  • Objective: To count the total number of resources classified as a dbo:City in the database.
PREFIX dbo: <http://dbpedia.org/ontology/>

SELECT (COUNT(?s) as ?numberOfCities)
WHERE {
  ?s a dbo:City .
}

4.2. Assessing Property Availability

  • Objective: To determine how many resources in the dataset have a specific property, which is useful for assessing data coverage and completeness.
  • Methodology: The following queries count entities with the dbo:populationTotal and dbo:spouse properties, respectively.
# Does the property dbo:populationTotal exist on any resource in the entire database?
PREFIX dbo: <http://dbpedia.org/ontology/>

SELECT (COUNT(?s) as ?entitiesWithPopulation)
WHERE {
  ?s dbo:populationTotal ?pop .
}
# List number of people with Spouse:
PREFIX dbo: <http://dbpedia.org/ontology/>

SELECT (COUNT(?s) as ?peopleWithSpouse)
WHERE {
  ?s dbo:spouse ?spouse .
}

4.3. Inspecting a Sample Entity from a Category

  • Objective: To retrieve all data for a single, arbitrary member of a class (dbo:City).
  • Methodology: A subquery with LIMIT 1 is used to efficiently select one city, whose properties are then fully retrieved by the main query.
# This query asks: "Find me one thing that is a city, and show me all of its properties."
SELECT ?p ?o
WHERE {
  # Sub-query to grab just one city URI
  {
    SELECT ?city WHERE { ?city a <http://dbpedia.org/ontology/City> . } LIMIT 1
  }

  # Main query to inspect that one city
  ?city ?p ?o .
}

4.4. Finding Cities in a Specific Country (India)

  • Objective: To list the 10 most populous cities in India, along with their state names.
  • Methodology: This query demonstrates a two-step traversal: it finds entities of type dbo:City, gets their dbo:state, and then checks if that state's dbo:country is India. The VALUES clause handles multiple possible URIs for India.
PREFIX dbo: <http://dbpedia.org/ontology/>
PREFIX dbr: <http://dbpedia.org/resource/>
PREFIX dbr-hi: <http://hi.dbpedia.org/resource/>
PREFIX foaf: <http://xmlns.com/foaf/0.1/>

SELECT ?cityName ?stateName ?population
WHERE {
  VALUES ?indiaURI { dbr:भारत dbr-hi:भारत } # Handle both possible URIs for India
  # Step 1: Find a city and its state
  ?city a dbo:City ;
        dbo:state ?stateURI ;
        dbo:populationTotal ?population ;
        foaf:name ?cityName .

  # Step 2: Check if that state is in India
  ?stateURI dbo:country ?indiaURI .

  # Optional: Get the name of the state too
  OPTIONAL {
    ?stateURI foaf:name ?stateName .
  }
}
ORDER BY DESC(?population)
LIMIT 10

4.5. Ranking by a Numeric Property

  • Objective: To find the 10 most populous cities in the entire Hindi DBpedia dataset, regardless of country.
  • Methodology: This query is similar to the one above but removes the constraints on state and country to create a global leaderboard.
# This query finds the largest cities in the entire dataset so even non indian cities that are present in our hindi graph
PREFIX dbo: <http://dbpedia.org/ontology/>
PREFIX foaf: <http://xmlns.com/foaf/0.1/>

SELECT ?cityName ?population
WHERE {
  ?city a dbo:City ;
        dbo:populationTotal ?population ;
        foaf:name ?cityName .
  # The dbo:country and dbo:state constraints are removed.
}
ORDER BY DESC(?population)
LIMIT 10

5. Advanced Queries and Data Types

This final section showcases more complex queries involving multiple relationships and specific data types like dates and geographic coordinates.

5.1. Finding Co-occurrence: Actors, Movies, and Directors

  • Objective: To find movies starring Amitabh Bachchan and list their directors.
  • Methodology: The query identifies movies where Amitabh Bachchan is the object of the dbo:starring predicate. For each movie, it retrieves its name and optionally retrieves the name of the person linked by the dbo:director property.
PREFIX dbo: <http://dbpedia.org/ontology/>
PREFIX dbr-hi: <http://hi.dbpedia.org/resource/>
PREFIX foaf: <http://xmlns.com/foaf/0.1/>

SELECT ?movieName ?directorName
WHERE {
  # 1. Find a movie where Amitabh Bachchan is in the starring cast
  ?movie dbo:starring dbr-hi:अमिताभ_बच्चन .

  # 2. Get the movie's name
  ?movie foaf:name ?movieName .

  # 3. OPTIONALLY find the director and their name
  OPTIONAL {
    ?movie dbo:director ?director .
    ?director foaf:name ?directorName .
  }
}
LIMIT 20

5.2. Filtering by Date

  • Objective: To find all people in the dataset born after January 1, 1950.
  • Methodology: This query retrieves all persons with a dbo:birthDate and uses a FILTER to compare the date value, which is explicitly cast as an xsd:date for correct comparison.
PREFIX dbo: <http://dbpedia.org/ontology/>
PREFIX foaf: <http://xmlns.com/foaf/0.1/>
PREFIX xsd: <http://www.w3.org/2001/XMLSchema#>

SELECT ?personName ?birthDate
WHERE {
  ?person a dbo:Person ;
          dbo:birthDate ?birthDate ;
          foaf:name ?personName .

  # The FILTER to check the date
  FILTER (?birthDate > "1950-01-01"^^xsd:date)
}
ORDER BY ?birthDate # Order by birth date, oldest first
LIMIT 50

5.3. Filtering by Text Pattern (Regex)

  • Objective: To find all universities or colleges whose names contain the Hindi word 'महाविद्यालय' (Mahavidyalaya/College).
  • Methodology: The query uses a UNION to select resources that are either a dbo:University or a dbo:EducationalInstitution. It then applies a REGEX filter on the name to find the specific pattern.
PREFIX dbo: <http://dbpedia.org/ontology/>
PREFIX foaf: <http://xmlns.com/foaf/0.1/>

SELECT ?institutionName
WHERE {
  # Find things that are EITHER a University OR an EducationalInstitution
  { ?institution a dbo:University }

  UNION

  { ?institution a dbo:EducationalInstitution }
  ?institution foaf:name ?institutionName .
  # The REGEX filter to find the specific word in the name
  FILTER(REGEX(?institutionName, "महाविद्यालय"))
}
LIMIT 50

5.4. Retrieving Geographic Coordinates

  • Objective: To find the geographic coordinates (latitude and longitude) for a sample of 10 cities.
  • Methodology: This query selects resources of type dbo:City and retrieves the values associated with the geo:lat and geo:long properties.
PREFIX dbo: <http://dbpedia.org/ontology/>
PREFIX foaf: <http://xmlns.com/foaf/0.1/>
PREFIX geo: <http://www.w3.org/2003/01/geo/wgs84_pos#>

SELECT ?cityName ?latitude ?longitude
WHERE {
  ?city a dbo:City ;
        foaf:name ?cityName ;
        geo:lat ?latitude ;
        geo:long ?longitude .
}
LIMIT 10

6. Conclusion

The queries presented in this document successfully demonstrate the capabilities of the Hindi DBpedia SPARQL endpoint. The endpoint can perform:

  • Entity Resolution: Finding resources based on literal values.
  • Property Traversal: Following single and multiple links between resources.
  • Data Aggregation: Counting and summarizing data across entire classes.
  • Complex Filtering: Filtering results based on location, dates, and text patterns (Regex).
  • Data Type Handling: Correctly processing specific data types such as dates and coordinates.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment