Skip to content

Instantly share code, notes, and snippets.

@myersjustinc
Last active October 5, 2015 16:38
Show Gist options
  • Select an option

  • Save myersjustinc/bf7ea7b60ac3f19ad7a3 to your computer and use it in GitHub Desktop.

Select an option

Save myersjustinc/bf7ea7b60ac3f19ad7a3 to your computer and use it in GitHub Desktop.
Computation and Journalism Symposium 2015 notes

Keynote: Lada Adamic, Facebook

Political news exposure

  • Pew: Majority of millennials, Generation X get political news from Facebook (for Baby Boomers: local TV)

  • Does this lead to a filter bubble or echo chamber?

  • Why does Facebook want to do this research?

    • Important to do research about how people use Facebook

    • Researchers all had studied content diversity and political discourse online before joining

    • Didn't set out to prove anything---just see what the state of things is

  • Twitter accounts and blogs that specialize in politics tend to follow others in a very polarizing pattern.

  • Expect potentially greater diversity on Facebook because friends are selected for variety of other reasons

    • You're less likely to share any given link from a weak tie than from a strong tie, but since you have so many more weak ties than strong, more of your shared links overall tend to come from weak ties.

    • Social signal---who shared this and how you otherwise know them---helps diminish reader's preference for/against the source of that content.

  • Factors also potentially could lower diversity on Facebook.

    • Diversity of friends who would share content to which you might be exposed could be low and lead to an echo chamber.

    • Algorithmic ranking could downrank cross-cutting content and lead to a filter bubble.

    • People could choose not to click on cross-cutting content and end up selectively exposed to certain content.

  • Is exposure diverse, or are there separate liberal and conservative echo chambers?

    • Used de-identified data from 10.1 million active U.S. users (25 percent) who self-report ideological affiliation, and > 7 million links shared

      • Mapped free-text affiliations on a scale (-2 (very liberal) to +2 (very conservative))

        • Left about 60 percent of affiliations unmapped
      • Limited, non-representative population

    • Identified hard news by training SVM classifier using strings such as "politi" for hard news and "entertain" for soft news

      • Classifier had 97.1 percent accuracy
    • Measured alignment of article by fraction of ideological affiliation of sharers (-1 (all liberal) to +1 (all conservative))

      • Bimodal distribution for hard news (echo chamber!)

      • More normal distribution for soft news (wider appeal!)

      • Validated against others' work on perceived alignment of news sources

  • How diverse is all content being shared on Facebook?

    • Slightly more conservative content is shared overall than liberal content.

      • Picking content at random, then, a liberal user would be slightly more likely to see cross-cutting content than a conservative user would.
  • How diverse are networks?

    • Moderates are more likely to have similar numbers of liberal and conservative friends.

    • Liberal and conservative networks are more likely to be homogeneous, but not starkly so.

  • How diverse is the content to which users can potentially be exposed, taking their friend networks into account?

    • Liberal users would come across considerably less cross-cutting content at this point than conservative users would.

    • This measure decreases slightly for both liberal and conservative users when you look at the content to which they actually are exposed.

  • Which articles to people click on?

    • Amount of cross-cutting content drops further still, but not a ton---ends up being about 30 percent for conservative users and about 20 percent for liberal users.
  • Users with weaker affiliations are exposed to and select slightly more cross-cutting content than users with stronger affiliations are.

  • Dataset is available and is reproducible!

The interesting lives of information cascades

  • Several kinds of memes actually do circle the globe.

  • How much social proof do you need before you share something or take some action?

    • Religious or political statements tend to need more social proof, which is defined by the proportion of your friends doing it first before you are willing to do it yourself.

    • Some people are more susceptible than others.

    • Combining several factors, it looks like users are willing to be convinced by friends spreading the meme up to a point---after that, those who haven't already spread it will tend to continue to resist.

  • Rumors tend to evolve over time and recur, but their popularity is bursty.

    • Parodies sometimes can be even more popular than the original rumors.
  • Why do memes persist?

    • If all memes were just as good and all users interacted at the same rates, there would be a short spike and no more.

    • If all memes were just as good and different users visited Facebook more or less frequently, there would be a wider, shorter bump, but still no more.

    • The variation of users' Facebook usage habits combined with the variation of memes' mutations' quality (and variants' quality) is what leads them to persist.

    • http://arxiv.org/abs/1402.6792

  • More research:

Panel: News Commenting, Moderation, and Community Systems

  • Participants:

    • Moderator: Andrew Losowsky, The Coral Project
    • Jenny Stromer-Galley, Syracuse University
    • Justin Cheng, Stanford University
    • Bassey Etim, New York Times

Bassey Etim

  • Etim's background: At Times for about seven years, community editor for three or four years of that

  • Why do users comment?

    Among those who responded that comments are extremely important or somewhat important to them:

    • Comments add diverse views (40 percent)
    • Comments are insightful and/or complement the article (36 percent)
    • Comments are of variable quality (5 percent)
    • Comments give readers freedom to express their opinion (4 percent)
  • Have to understand what you're trying to do and why you're trying to do it in order for a community (comments section or otherwise) to succeed

    • Prospective community members need to see some reflection of whatever it is they're seeking; if all they see is "pathologically angry" commenters, they're likely to be driven away (unless that's what they're after).
  • When Etim brings in a prospective moderator, he tries to "confuse the hell out of them".

    • Present with difficult comment and ask why to approve/disapprove
    • Ask about topics being covered
  • "We're failing every single day" we don't incorporate reader perspectives into our coverage.

Justin Cheng

  • Troll comments attract more troll comments, and things go off the rails really quickly.

  • How do we define trolling?

    • Can measure by post deletions and user bans

    • Literature defines trolling several ways that are both more specific and more ways.

    • For this work: Trolls are users who are eventually banned, and non-trolls are users who are similarly active but never banned.

  • How differently do trolls act?

    • Cosine distance between trolls' posts and previous posts is about 9 percent

    • Trolls attract twice as many responses

    • Trolls post more per thread

  • Do trolls change over time?

    • Unfairly deleting users' posts causes them to write worse content in the future.

    • About 1 in 4 trolls has a decreasing rate of post deletions in the second half of their life.

  • How do we predict troll-like behavior?

    • Can we predict whether a user will get banned in the future, based on their first 10 posts?

      • Yes---with similar accuracy to human moderators.
  • Can we implement systems where algorithms help assist human moderators?

    • Negative feedback makes users worse, whether those are actual responses or simple downvotes.

      • How can we discourage negative feedback?

        • Disqus moved from showing both upvotes and downvotes to showing a sum like Reddit does.

Jenny Stromer-Galley

  • Why do newsrooms allow comments in the first place?

    • Marketing/brand growth/followership

      • Community for the purpose of consumption
    • Readers as consumers

      • Allow for sharing of reactions and responses
    • Citizen journalism: readers as co-producers

      • Readers have useful and important ideas

      • Community for the sake of community

  • Why do people post and read comments?

    • Issue salience: They care about the topic and want to share their opinions

    • Hear others' opinions

    • Gain information

    • Examine one's own opinions

    • Vent

    • Sociability/community

    • Democratic empowerment

  • The problem with comments: Incivility

    • Recent study finds 20 percent of comments are uncivil

      • Name calling, pejoratives, vulgarity
      • Coe, Kenski, Rains, 2014
    • Uncivil commenters are infrequent commenters

    • Topic matters: Hard news is more uncivil than sports/lifestyle/health

    • Incivility referenced as why people stop participating or avoid it altogether

  • What promotes good comments:

    • Public identification (email, photos)
    • Moderation (automatic filtering, banning)
    • Interaction between staff and public
    • Shorter threads
  • What functions do comments serve in society?

    • Knowledge

      • Learn one's own views
      • Strengthen and stabilize people's views
    • Public opinion

      • Strengthen appreciation for other side of issue
  • What can we analyze comments for:

    • Discourse Quality Index

      • Measure rationality of argument

      • Considers:

        • Justification
        • Content of justification
        • Respect
        • Counterarguments
        • Constructive politics
        • Narrativity
    • Problems with algorithms for DQI

      • Misses other aspects of interaction

        • Rationality isn't the point
        • Community, play, storytelling
      • Loses community of the interaction

  • Advancing commenting

    • Topic, sentiment, position

      • How do we do this accurately?
    • Public opinion

    • Considering influencers

    • Cohesion, community

      • Help news staff interact with commenters
    • Automated agents to seed/foster good talk

      • Military funding research in this

Computational Approaches to Media Bias

Gender Discrimination by Audiences of Online News

  • Link: http://cj2015.brown.columbia.edu/papers/gender-discrimination.pdf

  • Authors:

    • J. Nathan Matias, Microsoft Research
    • Hanna Wallach, Microsoft Research
  • Research question: Is there a difference between how content by men and women is shared by others?

    • Research on differences in content sharing tends to focus on political differences
  • Women are a minority in U.S. newsrooms, opinion writing

    • More generally, treated differently in news articles
  • Efforts exist to change women's representation in news industry, media portrayal

  • Data collection: A year of three UK newspapers plus social media responses

    • Grouped by section and byline gender for subset of 156,523 articles that were clearly male or clearly female (by automated means)
  • Modeling approach: Random intercepts Poisson model

    • Byline level: Gender, number of articles
    • Article level: Title length, weekday, section
    • Interactions: Gender x Section
  • Selected findings for single-bylined stories

    • Guardian had 30 percent female-bylined stories, shared about on par with men's
    • Daily Mail had 20 percent female-bylined stories, but shared almost twice as much as men's
    • Telegraph had 19 percent female-bylined stories, but shared about 75 percent as much as men's
  • Limitations and future work

    • Topics and content
    • Periods of availability
    • Tenure at newspaper
    • Choosing what to control for
    • Appropriateness of models

Can Trending News Stories Create Coverage Bias? On the Impact of High Content Churn in Online News Media

  • Link: http://cj2015.brown.columbia.edu/papers/trending-news.pdf

  • Authors:

    • Abhijnan Chakraborty, Max Planck Institute for Software Systems
    • Saptarshi Ghosh, Max Planck Institute for Software Systems
    • Niloy Gangul, Indian Institute of Technology Kharagpur
    • Krishna Gummadi, Max Planck Institute for Software Systems
  • Background: News recommendations

    • News organizations produce way more information than any user can consume, so users need to rely on recommendations

    • What gets published in print is a recommendation, especially on the front page

    • Online recommendations are much more varied in source and in time scale

      • Sources: Individuals, experts, crowds, algorithms
      • Time scales: Future, now, past week
  • Question: Do different types of recommendations lead to different types of biases?

  • How do we measure bias?

    • Information diet
    • Composition of information produced or consumed
    • Compare diets of recommendations
  • For different sources of NYT recommendations (print, most-shared/emailed/commented, social), distributions of sections represented differs.

  • Need some amount of content churn in recommendations

    • Keep users coming back and more informed

    • Two time scales to keep in mind:

      • How often the recommendations are updated
      • The time frame from which stories are selected
    • But too much churn leads to missing late bloomers---stories that become popular only over the long term---despite being exposed to several times more content overall

      • Leads to bias against topics where stories have longer lives
    • Longer reads might be subject to time-of-day bias; science stories, for example, have a much more pronounced dip in the middle of the day.

  • Takeaways:

    • For news organizations:

      • Better understand and control biases of different recommendations being provided
      • Real-time systems and queuing theory might be useful
    • For news consumers:

      • Need to measure and manage information diets
      • Pay attention across multiple sites

Consumers and Suppliers: Attention asymmetries. A Case Study of Aljazeera’s News Coverage and Comments

  • Link: http://cj2015.brown.columbia.edu/papers/consumers-and-suppliers.pdf

  • Authors:

    • Sofiane Abbar, Qatar Computing Research Institute
    • Jisun An, Qatar Computing Research Institute
    • Haewoon Kwak, Qatar Computing Research Institute
    • Yacine Messaoui, Al-Jazeera
    • Javier Borge-Holthoefer, Qatar Computing Research Institute
  • Two models regarding criteria of news selection:

    • Trustee: Journalists decide what audiences should know
    • Market: Users read what they want
  • Some interaction between these models in practice---user behavior has time-lagged influence on later coverage decisions (others' previous research)

  • How should news organizations learn about and understand reader preferences?

    • Modeled as: Articles reflect journalists' attention, and comments reflect readers' attention.
  • Al Jazeera English comments

    • Two years
    • 22,000 articles
    • 2.3 million comments
    • 90,000 distinct users
    • 214 countries
  • Measuring which countries get covered versus which countries attract comments from which other countries

    • Few countries get most of the attention
  • Two patterns of comment attention (user attention):

    • Salient-event-driven: Something happened in or about one country, so users from many countries pay attention.

    • Regionalism-driven: Users from several neighboring countries pay attention to things happening in each other's countries.

  • Which countries get covered most by Al Jazeera journalists?

    • Mostly U.S. and Middle East
  • Which countries get commented on more often?

    • Still mostly U.S. and Middle East, but some
  • Certain countries commented on more than they're published:

    • U.S.
    • Israel
    • Egypt
    • Syria
    • Iraq
    • Russia
    • Turkey
    • Ukraine
    • Palestine
  • Who leads the news? Journalists or readers?

    • Comments cause (Granger-cause) articles of countries for which comments predominate.

    • Articles cause (Granger-cause) comments of countries for which articles predominate.

  • Future direction:

    • Generalize beyond Al Jazeera, perhaps with GDELT or social media
    • Build a platform to do this in real time

Computable Content: Structure, Context, Clarity, and Personalization

Improving the Comprehension of Numbers in the News

  • Link: http://cj2015.brown.columbia.edu/papers/numbers-in-news.pdf

  • Authors:

    • Pablo Barrio, Columbia University
    • Daniel Goldstein, Microsoft Research
    • Jake Hofman, Microsoft Research
  • XKCD describes the problem: http://xkcd.com/558/

  • Readers don't necessarily understand magnitude terms we use: http://bit.ly/trillionpoll

  • Try to augment readers' understanding: http://bit.ly/cut100million

  • Collected 370 perspectives from 80 Mechanical Turk workers on 67 NYT front-page article quotes

    • Had 1,800 different workers rate perspectives for perceived helpfulness (1 to 5 stars)
  • Tested 12 quotes for recall, estimation, error detection with and without sentences that provide additional perspective

    • Estimated with relative log error: abs(log(actual) - log(submitted)) / log(actual)

    • About half of respondents who received perspective sentences remembered the numbers they saw, compared with a third of those who didn't.

  • Are perspectives useful when estimating unknown quantities?

    • Provide a slider that updates the perspective sentence
    • Have the user guess the correct value
    • Perspectives reduced the bias and variance of responses
  • Do perspectives help in identifying potential errors?

    • Asked users whether figures looked plausible

    • Perspective helped about half of the time---inconclusive the rest of the time

      • So maybe you shouldn't do this all the time

      • Helped less when people's initial guesses were correct

DeScipher: A Text Simplification Tool for Science Journalism

  • Link: http://cj2015.brown.columbia.edu/papers/descipher.pdf

  • Authors:

    • Yea Seul Kim, University of Washington
    • Jessica Hullman, University of Washington
    • Eytan Adar, University of Michigan
  • Science journalism faces significant language barriers.

  • How do science journalists make scientific findings accessible?

    • Simplification strategies

      • Replace unfamiliar terms with more familiar terms
      • Define unfamiliar terms
      • Use short sentences
    • Humanize the scientist or those affected with anecdotes and stories

    • Provide relevant background information

    • Use visualizations and figures

  • Focused on first simplification strategy (replace unfamiliar terms)

    • Built automated tool (Sublime plugin?) to help detect jargon and suggest alternatives
  • Design motivation: keep journalist in loop

    • Journalist's input should be augmented, not replaced.
  • Data: PLoS requires authors to submit both abstracts (more technical) and author summaries (more accessible)

    • Obtained almost 16,000 pairs of abstracts and author summaries
  • Pipeline design:

    • Identify candidate simplification rules

      • Extract content words (remove stop words)
      • Generate pairs of words that might go together
    • Filter simplifications

      • Retain synonyms/hypernyms
      • Calculate complexity to ensure (complex, simple) pair
      • Confirm grammatical interchangeability
    • Apply simplifications

      • Is a word being used precisely for its complex meaning?
      • Does the word mean what we think it means?
    • Display multiple ranked simplifications to journalist

      • Looks like a Sublime menu
  • Evaluation via Mechanical Turk: Which is more complex? How helpful is it?

  • Future evaluation:

    • Study journalists' processes
    • Identify how DeScipher might fit into those processes
    • Find other uses/features/functions that might be useful

Editorial Aspects of Reporting into Structured Narratives

  • Link: N/A

  • Authors:

    • David Caswell, Structured Stories/University of Missouri
    • Bill Adair, Duke University
    • Frank Russell, University of Missouri
  • Does the article make sense as the primary unit of news, either for newsrooms or for news consumers?

  • What are the constituent semantic parts of news that can be disassembled and reassembled?

    • Hypothesis: These are events and stories.

      • Event: Instance of specific activity with specific participants in specific roles and bounded in space and time

      • Story/narrative: Knowledge-representation structure for storing and optimally organizing event information.

  • Structured journalism: Store individual components in a central repository from which new presentations can be generated as needed.

Putting news in context, automatically

  • Link: http://cj2015.brown.columbia.edu/papers/news-in-context.pdf

  • Authors:

    • Larry Birnbaum, Northwestern University
    • Miriam Boon, Northwestern University
    • Scott Bradley, Northwestern University
    • Jennifer Wilson, Northwestern University
  • Providing context should be part of the journalistic standard of care, but only 12 percent of surveyed journalistic codes of ethics (n ~= 250) mentioned anything resembling context---why?

    • No longer limited by available publication space and time, but by user interest/time and by scalability
  • Contextual systems should provide information that is somewhat similar (i.e., relevant) to what the user is looking at, but it also should be somewhat different (i.e., novel) in order to provide value.

    • Provide, for example, additional actors, additional figures or additional quotes.
  • Built several systems to do this in different ways to different kinds of content, and realized the architecture had a lot of similarities

    • Abstracted this out into an open source toolkit (The News Context Project) to use in future projects

      • Still requires some developer effort to apply, though---how can this be made user-configurable?

        • Interface could, for example, represent tradeoffs between types of content desired

Designing for Personalized Article Content

  • Link: http://cj2015.brown.columbia.edu/papers/personalized-content.pdf

  • Authors:

    • Carolyn Gearig, University of Michigan
    • Eytan Adar, University of Michigan
    • Jessica Hullman, University of Washington
  • Looking at personalized content

    • Defined as an automated change to the set of facts in an article's content based on properties of the reader

    • As opposed to personalized feeds of content that doesn't change

    • Can include:

      • Adding facts
      • Removing facts
      • Providing context/emphasis
    • How does this work, and how can we facilitate it?

  • Looking into risks, benefits and guidelines

  • Surveyed 22 journalism professionals for perceived risks and benefits

    • Most surveyed professionals involved with graphics production

    • Most surveyed professionals defined personalization differently than researchers did

  • Benefits

    • Increase reader engagement and learning
    • Allow broad publications to appeal to specific audiences
    • Allow journalists to write one article that applies to many readers
  • Risks

    • Could overstep reader privacy---most common concern journalists had
    • Could misrepresent content
  • How would journalists personalize content?

    • Most common responses: location, internet habits, age
  • Guidelines

    • Personalize with function in mind

      • Decision to personalize should be deliberate and not overused
    • Interactivity != personalization

      • At very least, personalization should happen automatically and not require interaction
    • Consider inference quality at all levels

    • Identify failure and fail gracefully

    • Privacy is a crucial concern

    • Identify bias that could be encountered

    • Provide reader control

      • Readers should know how it's being personalized and correct any assumptions that have been made

      • Readers should be able to opt out if desired

    • Journalists' workflows are unique

  • Implementation

    • Unlike lab's other projects, this aims to keep journalists more in the loop---not automating as many decisions

    • Many survey participants believe this kind of work is too complicated for their organizations

    • Aiming to simplify this with PersaLog:

      • Inference engine to determine reader characteristics
      • Domain-specific language to describe personalization logic
      • Public-facing visualization components
    • Can shift where computation is performed

      • On server, which is more powerful but introduces additional privacy concerns

      • On client, which is less powerful but more under reader's direct control

    • Focuses on logic---data work still needs to be done to give you something to display

  • Still in development---sign up for announce list: http://goo.gl/D22c3t

Progression Beyond the Mean: Journalists Tackle Complex Modeling

Machine Learning in the Wild

  • Speaker: Janet Roberts, Reuters

  • The Echo Chamber

  • Started with pitch from SCOTUS beat reporter

    • Let's look at the subset of lawyers who are at SCOTUS all the time.

    • That story's been told, though.

  • Out of 10,000 appeals submitted each year, only 75 get heard.

    • This is the most important step in the process---how does an appellant's choice of lawyer affect his or her odds?

    • 17,000 lawyers involved in these appeals---66 of them accounted for 43 percent of the cases that ended up being heard.

      • Is there an access-to-justice issue here?
  • Had freelancers code petitions by case type and petitioner type over case of three months

    • Heard at NICAR about Northwestern students categorize Congressional Record by topic
  • Had someone on Reuters' R&D team apply LDA to this

    • Couldn't use all detected topics, but could use the best ones from each petition

    • Certain topics were more reliably identified than others

  • Combination of machine leaning and human checking ended up really powerful

    • Ended up customizing stop word list to help clean things up

    • Had to experiment with number of topics to get a not-too-big, not-too-small, not-too-boring, not-too-weird set

Medicare

Creating Surgeon Scorecard

  • Speaker: Olga Pierce, ProPublica

  • Problem: Patients don't know where to go for the safest care, and patient harm is the third-leading cause of death in America.

  • Perfect data is unavailable---all that was available was Medicare billing data, which doesn't include complete clinical records (just basic data).

    • Even this can be unreliable, so ProPublica only used the most reliable data points: deaths in the hospital and readmissions within 30 days for surgery-related complications.

    • Risk adjustment poses other problems, since some cases are more complex than others---so they chose procedures where this was less of an issue: low-risk, elective procedures, excluding ER admissions, transfers and unusual diagnoses.

  • Aimed to produce one relatively easy-to-understand measure, where the units are people.

  • Expected to analyze hospitals, but it turns out there's a lot of variation among surgeons in the same facility.

    • Surgeon IDs in the data were opaque---couldn't tell who individual surgeons were until someone else sued for the crosswalk.
  • Analysis:

    • Identified surgeons who performed these procedures on Medicare patients

    • Identified cases where complications took place

    • Ended up with raw complication rate

    • Adjust for risk with mixed-effects model:

      • Adjust for patient health and age
      • Isolate hospital effect
      • Determine surgeon effect
    • End up with adjusted complication rate and confidence interval for each surgeon/procedure combination

  • Reran model five times to get five different confidence intervals; this let them display a gradient with that interval to convey different levels of confidence.

    • Important since confidence levels add transparency and can help address potential misclassifications.
  • Used different model for geographic search, so users saw nearby hospitals ordered by quality of hospital/surgeon combinations, even though the raw values were never shown to them

Data Journalism in the Classroom

State of Data Journalism Education

  • Paper: N/A

  • Speaker: Cheryl Phillips, Stanford University

  • Data journalism education started outside journalism schools

    • Really started in the 1970s with Phil Meyer's work and others' imitations

    • Journalists started teaching each other and organizing conferences

    • Wasn't until the mid-1990s until university journalism programs officially started teaching data analysis for stories, and even then it was rare

  • Looked at 113 AEJMC-accredited journalism programs

    • Also collected about 55 syllabi and interviewed instructors
  • Level of instruction still has room to grow

    • 54 programs do not offer any data journalism classes at all
    • 27 programs only offer one class
    • 14 offer two classes
    • 18 offer three or more classes
  • Vast majority don't offer data visualization courses

  • Key concepts from data journalism syllabi, in order:

    • Journalistic/critical thinking
    • Spreadsheets
    • Relational databases
    • Design concepts
    • Statistics basics
    • Programming concepts
  • Students found out about these classes late in their programs and didn't have room to explore after taking these classes if they had the flexibility to take them in the first place

  • Few core textbooks used in courses---only about 10 core books being taught

  • Fewer than 20 journalism schools teach any kind of programming beyond HTML

  • Most instructors are former data journalists or current data journalists

    • No pervasive effort to try to change nature of programs to include data
  • Need more than skills---need people with experience in pedagogy to help instruct and inspire

"Flipping" for Journalism Tech Education

  • Paper: http://cj2015.brown.columbia.edu/papers/flipping.pdf

  • Speaker: Susan McGregor, Columbia University

  • "Should journalists learn to code?" debate misses the nuance of what we need to teach and why

    • Not necessarily CS, but at least scripting

      • Independence
      • Critical evaluation skills
      • Flexibility
      • Literacy
  • Challenges in teaching technology:

    • Self-concept/self-esteem/stereotypes

      • McGregor starts each class with at least the intro from "Program or Be Programmed: Ten Commands for a Digital Age" by Douglas Rushkoff

        • People feel like they can learn to drive if they want to, even if they don't always use it

        • People encounter and observe driving in their daily lives

        • People see a wide variety of people driving every day

      • Programming languages are languages---and less ambiguous ones than natural language

        • Journalism students already are good with language
    • Abstractness of concepts

      • "Translate"/rewrite a story in code

        • "In the course of one 300-word story, you can find examples of every kind of data structure you're going to need."
    • Scalability of instruction

      • Flipped instruction!

        • Students can better direct their attention/questions in class

        • Students can review material as needed, minimizing individual questions

        • Ability to review reduces anxiety

        • Production time is minimal

        • Questions unaddressed in class can be integrated via additional videos

      • Don't edit out mistakes---important to show how you recover from errors

  • Important to frame concepts/motivations more accessibly---program solving is important and understandable

    • Basic math instruction gets this to some extent
    • "Grounding things in concrete experiences" is important
  • Have to be able to think through concepts as a novice would

Teaching Coding in Journalism Schools: Considerations for a Secure Technological Infrastructure

  • Paper: http://cj2015.brown.columbia.edu/papers/teaching-coding.pdf

  • Speaker: Meredith Broussard, New York University

  • Vast range of possible things that could be taught in data journalism

    • Focus on concepts/literacy
    • Lots of math anxiety
    • Not necessarily a lot of experience with multiple platforms
  • Leveraging university's technological resources can be trickier for instructor, but it's important for privacy and accessibility

    • Students should be able to build on what instructor provides rather than having to jump directly into interacting with outside vendors.

    • Students learning about data journalism deserve the same privacy/comfort as students learning about calculus

  • Need to allow students to make use of university computing resources because they don't necessarily have financial means to acquire their own

    • Accessibility to educational opportunity is extremely important

    • University IT people don't like the possibility of students making administrative errors on their machines

      • Broussard testing out approach with NYU IT-managed AWS resources

Opening the Process

No Data, No Computation, No Replication or Re-Use: The Utility of Data Management and Preservation Practices for Computational Journalism

  • Paper: http://cj2015.brown.columbia.edu/papers/data-management.pdf

  • Speakers:

    • Kris Kasianovitz, Stanford University
    • Regina Roberts, Stanford University
  • Many data lifecycles omit librarians/archives

  • Preservation is at least as important as archiving

    • Suggested formats for archiving:

      • Image: JPEG, JPEG 2000, PNG, TIFF
      • Text: Plain text, HTML, XML, PDF/A
      • Audio: AIFF, WAV
      • Containers: tar, gzip, zip
      • Databases: XML or CSV
    • Resource list and demo materials: http://purl.stanford.edu/zh188pk0040

  • Data includes both unstructured and structured data

    • Software (with version information)
    • Algorithms used for analysis
    • Files used for analysis
    • Metadata
    • Codebooks
    • Licenses and rights statements
    • Detailed READMEs
  • Remember, GitHub does not guarantee any sort of long-term archiving. Keep things in institutional repositories as well.

  • What restrictions are there on your data?

    • HIPAA
    • FERPA
    • PII
    • Financial information
    • Proprietary
    • Fee-for-access
  • Clarify what you can:

    • Share with other researchers

    • Share with researchers only within your institution/organization

    • Share only variables, codes, methods and contact information from data source

    • Share extracts, not entire dataset, including text data mining outputs

    • Put in a dark archie

  • Case study: Local government and open data

  • Case study: Text and data mining proprietary data

    • Sometimes most appropriate corpora are restricted by copyright and/or institutional license agreements

      • Some publishers/aggregators allow access to APIs for metadata and texts that are no longer under copyright

      • Need exists for new or renegotiated license agreements

        • Librarians/archivists often can help you with these negotiations since they've fought these battles before
    • Secure virtual enclaves, such as ICPSR's: https://www.icpsr.umich.edu/icpsrweb/content/ICPSR/access/restricted/enclave.html

  • Recommendations:

    • Develop new (or follow existing) data management plan---think through data lifecycle

    • Document your methods

    • Document your software, algorithms and/or scripts

    • Seek secure data storage

    • Work with data librarians or archivists in order to deposit your data and metadata into an institutional repository (local or other)

    • Use licenses to clearly indicate how others may reuse your data

The Gamma: Programming tools for transparent data journalism

  • Paper: http://cj2015.brown.columbia.edu/papers/gamma.pdf

  • Speaker: Tomas Petricek

  • http://thegamma.net/

  • News organizations produce a lot of static charts

    • Data behind those charts isn't often available
  • Data work should be reproducible

  • A data-driven article can be a literal program

  • The Gamma includes source code and customizability for all visualizations in its articles

    • The article is just one view of the underlying program
  • Editing tool should integrate with many data sources and make programming easy

  • Interfaces are important to help user navigate multidimensional data sources

    • Autocomplete can be powerful for this
  • Summary

    • Encourage active information literacy
    • Articles as programs are power
    • Make future usable interfaces to programming

The Quest to Automate Fact-Checking

  • Paper: http://cj2015.brown.columbia.edu/papers/automate-fact-checking.pdf

  • Speakers:

    • Naeemul Hassan, University of Texas at Arlington
    • Bill Adair, Duke University
    • James Hamilton, Stanford University
    • Chengkai Li, University of Texas at Arlington
    • Mark Tremayne, University of Texas at Arlington
    • Jun Yang, Duke University
    • Cong Yu, Google Research
  • Technical hurdles for popup fact-checking

    • Most fact-checks are published in old blog formats such as WordPress

    • Most are not in a structured-journalism format that would allow for more flexibility and easier matching.

    • Current publishers have inconsistent relational coding (if any) for video and campaign ads.

    • Vagueness and natural language variation can make it harder to relate a new fact to a previous fact check.

  • Next steps

    • Develop apps to detect political messages in video and audio
    • Build ClaimBuster, which detects factual claims
  • ClaimBuster

    • Demo: http://idir.uta.edu/claimbuster

    • Take transcripts of presidential debates

    • Have humans annotate those transcripts

      • Can categorize sentences:

        • Important factual claims

          • Main target of fact checking

          • "We spend less on the military today than at any time in our history."

        • Unimportant factual claims

          • "I was in Iowa yesterday."
        • Sentences with no factual claims (e.g., opinions, questions)

          • "I will be tough on crime."
      • Had humans code a training set of sentences

    • Extract features with Alchemy and Python

      • Find most important features to extract---most important was a cardinal number
    • Apply a learning algorithm

      • Support vector machines performed best
  • Case study: first GOP debate in 2015

    • Near-real-time experiment using live transcript from closed-captioning data (from TextGrabber device)

    • Included 1,393 sentences

    • 71 percent of facts checked by CNN, factcheck.org and PolitiFact were included in first 18 percent of sentences

Ranking the Age of Algorithms and Curated News

  • Participants:

    • Anthony De Rosa, Circa
    • Catherine D'Ignazio, Emerson College
    • Mike Dewar, The New York Times
    • Moderator: Suman Deb Roy, betaworks
  • Ranking algorithms are making their way into more and more content systems, but the developers and entrepreneurs lack either tact or understanding of underlying behavior.

  • Dewar: Must determine who's problem we're trying to solve

    • Often developers are trying first and foremost to solve their own problems, not users' problems.

    • The problem definition has many implications for decisions being made further along in the process.

  • De Rosa: Trending topics are useful to help direct limited newsroom resources, and they're useful to help readers follow particular stories.

    • Users have relatively few options for the latter right now.
  • D'Ignazio: Developed geographic entity extraction system for news articles: http://cliff.mediameter.org/

  • Dewar: "I get frustrated when people think it's a kind of science--and not design."

  • De Rosa and Roy: Annoying push notifications are a leading reason people uninstall apps---be very intentional and very careful about when and how these are sent.

  • D'Ignazio: Need signals for decision-making that aren't just quantitative; do more ethnographic and qualitative research as well.

  • Dewar and Roy: Selecting algorithms and parameters is a lot of trial and error. There's no rule of thumb---more just seeing what ultimately works best.

  • D'Ignazio and Dewar: Interfaces and presentation should be more transparent in order to help improve users' literacy about data.

Visualization for Story Finding and Telling

RevEx: Visual Investigative Journalism with A Million Healthcare Reviews

  • Participants:

    • Cristian Felix, New York University
    • Anshul Vikram Pandey, New York University
    • Enrico Bertini, New York University
    • Charles Ornstein, ProPublica
    • Scott Klein, ProPublica
  • http://nyuvis.github.io/revex

  • ProPublica came to NYU researchers with 1.3 million Yelp reviews of health care providers---how should they sift through this?

  • Journalistic intent is important

    • Need to spot doctors who are bad actors
  • Provide the journalists with ability to make exploratory searches with accompanying statistical analyses

    • How does this sub-sample differ from another sub-sample or the overall sample?
  • Noise is a big problem---quality of reviews differs dramatically.

  • RevEx allows for both faceted and full-text searches and also provides suggestions for other search terms.

  • Data isn't a proxy for the truth, of course.

    • ProPublica followed up in person on findings it made.
  • Hope to generalize to use with different datasets in the future

Lenses: An Open-Source Tool for Instant Data Visualizations

  • Participants:

    • Amy Chen
    • Kareem Amin
    • Helen Carey
    • Nivvedan Senthamil Selvan
  • Project came from News Corporation and was worked on with NYU and Columbia

  • Community platform for data manipulation/visualization

    • Connect popular open datasets

    • Drag-and-drop interface

    • Define "lenses"---specific views/transformations

      • Chain together series of components for input, transformation, visualization

      • Data flows through a lens pipeline

    • Open source and easily extensible

  • Core philosophy

    • Replicability and transparency

      • History preserving

      • Each visualization can be traced back through the transformations to the source data

      • Data can be inspected before and after each transformation

      • Can be easily shared and checked for accuracy

    • Extensibility and open source

    • Community and reusability

      • Visualizations and transformations are saved and searchable

      • Aim to provide a GitHub-style model for data visualizations

      • Let others look at how peers have used data and fork/extend that work

  • Hope to collaborate with editors, subject matter experts, etc.

Tweet Location Detection

  • Participants:

    • Bahareh R. Heravi, National University of Ireland
    • Ihab Salawdeh, National University of Ireland
  • Many Twitter visualizations are about large events and large populations, but looking at Ireland, the population is considerably smaller (even for the Ireland/Scotland Euro 2016 group qualifying match).

  • Many visualizations use only geotagged tweets, but only 1 to 5 percent of tweets are geotagged.

    • This might not be enough for some applications.
  • Services exist to add geotagging, but the starting price is on the order of USD 10,000 per year---prohibitive for many news organizations.

  • Location information sources

    • Geotagged tweets

    • User-specified profile locations

      • Explicitly specified location
      • Time zone
      • Language
    • Entity extraction/natural language processing

      • Extract locations mentioned or locations of organizations mentioned
    • Social network analysis

      • Not including this yet

        • Not that specific
        • More computationally intensive
  • In general, method is to prefer the most accurate data and

  • Out of 22,957 tweets during the match, 1,055 were geotagged, but they managed to georeference 16,008 total using this technique.

Keynote: Chris Wiggins, The New York Times

  • Why does The Times have a data science team?

    • Biology changed for the better by becoming data-driven, as an example.

      • (Wiggins has a background in computational biology, among other fields.)
    • How do we combine hacking skills, substantive expertise, and math and statistics knowledge?

    • Wiggins had a sabbatical and decided to take it at The Times.

    • This team is not on the editorial side. ("Firmly in state.")

    • The Web isn't just an "online presence"---it's a microscope through which to analyze all kinds of things, including your audience, and it's an experimental tool to help optimize your work.

  • What does the team do there?

    • Modeling

      • Predictive: What will happen?

        • How many people will subscribe? How many will stop?

        • How many papers should we print? Where should we send them?

        • Which pieces of content will users want to read?

      • Descriptive: What's happening?

        • How do different people interact with content?
      • Prescriptive: What should I do?

        • What should we post to Facebook and when, in order to get it to as many people as possible?
  • How can other large organizations start and integrate data science teams?

    Project stage Modeling type
    Explore Descriptive
    Predict Predictive
    Test
    Optimizing Prescriptive
    Reporting
    • Common requirements for culture change, per U.S. Air Force:

      • People

        • Having a new mindset is much more important than having a new toolset.
      • Ideas

        • Skills: Data engineering, data science, data visualization, data product, data multiliteracies, data embeds

          • Data engineering

          • Data science

          • Data visualization

          • Data product

          • Data multiliteracies

            • How can people think critically about data, either technically or rhetorically?
          • Data embeds

            • Work with others in the organization!
      • Things

        • Deliverables: Prototypes, APIs, impact roadmaps

Takeaways

  • Exploratory data analysis: Some NYU researchers built a system for ProPublica to use when sifting through a bunch of Yelp reviews for a story they ran with NPR on those reviews.

    They plan to extend it later to work with other kinds of text UGC.

  • Data preservation: We're already doing a lot of this well, but a couple of Stanford librarians described a number of things to make sure to preserve for data projects.

    They also reminded the room that GitHub alone does not count as an archival strategy. Being academic librarians, they suggested submitting as much as possible to some sort of institutional repository, whether that's an internal one or something like ICPSR or the Stanford Digital Repository.

  • Numerical literacy: Some researchers from Microsoft described their efforts to present and test different ways of putting numbers in context within news stories. (For example, a 400-foot ceiling on drone flights is about the height of a 40-story building.)

    Of course, this is a topic that's come up in xkcd before.

  • State of data journalism education: Cheryl Phillips of Stanford surveyed more than 100 accredited university journalism programs in the U.S. and found a serious lack of data journalism training.

    This isn't particularly surprising, but it is depressing.

  • Machine Learning: Mike Dewar of The New York Times pointed out that machine learning in particular is a very involved kind of design process, rather than a science per se.

    More generally, he says selecting the right algorithm and parameters for a project isn't the sort of thing that has good rules of thumb; it's a lot more of a trial-and-error process than people tend to think.

  • Fact checking: A paper on automated fact checking had a potentially useful taxonomy of types of sentences in a political speech or debate:

    • Important factual claims (main target of fact checking)

      For example: "We spend less on the military today than at any time in our history."

    • Unimportant factual claims

      For example: "I was in Iowa yesterday."

    • Sentences with no factual claims (such as opinions or questions)

      For example: "I will be tough on crime."

  • Book recommendation: Program or Be Programmed, by Douglas Rushkoff

    Describes programming as a kind of technical literacy that should be encouraged.

    The intro apparently makes a lot of comparisons to driving, since that's another way of operating a machine that some people (but not most) do professionally:

    • People feel like they can learn to drive if they want to, even if they don't always use it.

    • People encounter and observe driving in their daily lives.

    • People see a wide variety of people driving every day.

    Described in the context of introductions to technical concepts within journalism curricula.

  • Entity extraction: The Center for Civic Media at MIT has built an entity extraction system that's made specifically for how news articles tend to be written.

    It also includes a geographic parser, which could be interesting.

  • Trolls: A Stanford researcher had some relatively simple ways to define and quantify trolls and their behavior.

  • Recommendation problems: Some researchers in Germany and India looked at NYT recommendations with different time scales and how certain stories tend to be sort of late bloomers and the potential for coverage bias that results from that.

    If a story isn't doing well on your daily or weekly top-content lists, will you be discouraged from writing that kind of thing in the future? What if it turns out that same story actually is one of the month's top performers? Will you even notice?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment