Skip to content

Instantly share code, notes, and snippets.

View jaceklaskowski's full-sized avatar
:octocat:
Enjoying developer life...

Jacek Laskowski jaceklaskowski

:octocat:
Enjoying developer life...
View GitHub Profile
@jaceklaskowski
jaceklaskowski / dcos.md
Last active June 30, 2018 12:13
Introduction to DC/OS
@jaceklaskowski
jaceklaskowski / sparkathon-agenda.md
Last active October 10, 2017 20:28
Sparkathon in Warsaw - Development Activities
@jaceklaskowski
jaceklaskowski / scala-something.md
Created February 24, 2017 22:42
Scala SOMETHING

Notes

  • Combinators == building blocks
  • functions and higher order functions
  • composition
  • immutable == less moving parts to worry about
  • a routine job == a boilerplate == a boring stuff
  • a context (so the job varies)
  • abstracting away == happening behind the scenes == cutting down repetitive code == eliminating boilerplate
  • "The code becomes small, succinct, and more readable"
@jaceklaskowski
jaceklaskowski / parquet.md
Last active December 26, 2017 19:16
Parquet

Parquet

Introduction

  • Stores schema information along with the data
  • Columnar storage/file format
    • "reference file format on Hadoop HDFS"
    • "read-optimized view of data"
  • excellent for local file storage on HDFS (instead of external databases).
  • writing very large datasets to disk
@jaceklaskowski
jaceklaskowski / blockchain.md
Last active February 17, 2026 23:47
Blockchains, Cryptoeconomics, Ethereum, Litecoin, Bitcoin, IOTA
@jaceklaskowski
jaceklaskowski / spark-exercises.md
Last active June 26, 2022 12:04
Spark Exercises

Exercise 1

Union only those rows (from large table) with keys in left small table, i.e. union two dataframes together but only those with the key in my small table.

Exercise 2

Aggregation on an array of nested json = How to sum the quantities across all lines for a given order (which would give 1 + 3 = 4 for the below sample dataset):

{

IDEA:

  • breaks on demand given the number of exercises
  • break man who says we should have one
val wholeJsonRDD = sc.wholeTextFiles("input.json").map(_._2)
val mySchema = new StructType().add($"n".int)
wholeJsonRDD.toDF.withColumn("json", from_json($"value", mySchema)).show(truncate = false)

val jsonDF=spark.read.json("output.json")

Exercise

Develop a Spark standalone application (using IntelliJ IDEA) with Spark MLlib and LogisticRegression to classify emails.

Think about command line and what parameters you'd like to accept for various use cases.

TIP Use scopt

  1. libraryDependencies += "org.apache.spark" %% "spark-mllib" % "2.1.1"
@jaceklaskowski
jaceklaskowski / anatolyi.md
Created December 16, 2017 12:14
Anatolyi - Facebook Profiles
// Let's create a sample dataset with just a single line, i.e. facebook profile
val facebookProfile = "ActivitiesDescription:703 likes, 0 talking about this, 4 were here; Category:; Email:joe@pvhvac.com; Hours:Mon-Fri: 8:00 am - 5:00 pm; Likes:703; Link:https://www.facebook.com/pvhvac; Location:165 W Wieuca Rd NE, Ste 310, Atlanta, Georgia; Name:PV Heating & Air; NumberOfPictures:0; NumberOfReviews:26; Phone:(404) 798-9672; ShortDescription:We specialize in residential a/c, heating, indoor air quality & home performance.; Url:http://www.pvhvac.com; Visitors:4"
val fbs = Seq(facebookProfile).toDF("profile")

scala> fbs.show(truncate = false)
+----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
@jaceklaskowski
jaceklaskowski / person.md
Last active January 8, 2019 19:37
PersonSerde
  case class Person(id: Long, name: String)

  class PersonSerializer extends Serializer[Person] {
    override def configure(configs: util.Map[String, _], isKey: Boolean): Unit = {}

    override def serialize(topic: String, data: Person): Array[Byte] = {
      println(s">>> serialize($topic, $data)")
      s"${data.id},${data.name}".map(_.toByte).toArray
    }