中文字幕av专区_日韩电影在线播放_精品国产精品久久一区免费式_av在线免费观看网站

溫馨提示×

溫馨提示×

您好,登錄后才能下訂單哦!

密碼登錄×
登錄注冊×
其他方式登錄
點擊 登錄注冊 即表示同意《億速云用戶服務條款》

如何理解Receiver啟動以及啟動源碼分析

發布時間:2021-11-24 16:05:31 來源:億速云 閱讀:124 作者:柒染 欄目:云計算

今天就跟大家聊聊有關如何理解Receiver啟動以及啟動源碼分析,可能很多人都不太了解,為了讓大家更加了解,小編給大家總結了以下內容,希望大家根據這篇文章可以有所收獲。

為什么要Receiver?

Receiver不斷持續接收外部數據源的數據,并把數據匯報給Driver端,這樣我們每隔BatchDuration會把匯報數據生成不同的Job,來執行RDD的操作。

Receiver是隨著應用程序的啟動而啟動的。

Receiver和InputDStream是一一對應的。

RDD[Receiver]只有一個Partition,一個Receiver實例。

Spark Core并不知道RDD[Receiver]的特殊性,依然按照普通RDD對應的Job進行調度,就有可能在同樣一個Executor上啟動多個Receiver,會導致負載不均衡,會導致Receiver啟動失敗。

Receiver在Executor啟動的方案:

1,啟動不同Receiver采用RDD中不同Partiton的方式,不同的Partiton代表不同的Receiver,在執行層面就是不同的Task,在每個Task啟動時就啟動Receiver。

這種方式實現簡單巧妙,但是存在弊端啟動可能失敗,運行過程中Receiver失敗,會導致TaskRetry,如果3次失敗就會導致Job失敗,會導致整個Spark應用程序失敗。因為Receiver的故障,導致Job失敗,不能容錯。

2.第二種方式就是Spark Streaming采用的方式。

在ReceiverTacker的start方法中,先實例化Rpc消息通信體ReceiverTrackerEndpoint,再調用

launchReceivers方法。

/** Start the endpoint and receiver execution thread. */
def start(): Unit = synchronized {
  if (isTrackerStarted) {
    throw new SparkException("ReceiverTracker already started")
  }

  if (!receiverInputStreams.isEmpty) {
    endpoint = ssc.env.rpcEnv.setupEndpoint(
      "ReceiverTracker", new ReceiverTrackerEndpoint(ssc.env.rpcEnv))
    if (!skipReceiverLaunch) launchReceivers()
    logInfo("ReceiverTracker started")
    trackerState = Started
  }
}

在launchReceivers方法中,先對每一個ReceiverInputStream獲取到對應的一個Receiver,然后發送StartAllReceivers消息。Receiver對應一個數據來源。

/**
 * Get the receivers from the ReceiverInputDStreams, distributes them to the
 * worker nodes as a parallel collection, and runs them.
 */
private def launchReceivers(): Unit = {
  val receivers = receiverInputStreams.map(nis => {
    val rcvr = nis.getReceiver()
    rcvr.setReceiverId(nis.id)
    rcvr
  })

  runDummySparkJob()

  logInfo("Starting " + receivers.length + " receivers")
  endpoint.send(StartAllReceivers(receivers))
}

ReceiverTrackerEndpoint接收到StartAllReceivers消息后,先找到Receiver運行在哪些Executor上,然后調用startReceiver方法。

override def receive: PartialFunction[Any, Unit] = {
  // Local messages
  case StartAllReceivers(receivers) =>
    val scheduledLocations = schedulingPolicy.scheduleReceivers(receivers, getExecutors)
    for (receiver <- receivers) {
      val executors = scheduledLocations(receiver.streamId)
      updateReceiverScheduledExecutors(receiver.streamId, executors)
      receiverPreferredLocations(receiver.streamId) = receiver.preferredLocation
      startReceiver(receiver, executors)
    }

startReceiver方法在Driver層面自己指定了TaskLocation,而不用Spark Core來幫我們選擇TaskLocation。其有以下特點:終止Receiver不需要重啟Spark Job;第一次啟動Receiver,不會執行第二次;為了啟動Receiver而啟動了一個Spark作業,一個Spark作業啟動一個Receiver。每個Receiver啟動觸發一個Spark作業,而不是每個Receiver是在一個Spark作業的一個Task來啟動。當提交啟動Receiver的作業失敗時發送RestartReceiver消息,來重啟Receiver。

/**
 * Start a receiver along with its scheduled executors
 */
private def startReceiver(
    receiver: Receiver[_],
    scheduledLocations: Seq[TaskLocation]): Unit = {
  def shouldStartReceiver: Boolean = {
    // It's okay to start when trackerState is Initialized or Started
    !(isTrackerStopping || isTrackerStopped)
  }

  val receiverId = receiver.streamId
  if (!shouldStartReceiver) {
    onReceiverJobFinish(receiverId)
    return
  }

  val checkpointDirOption = Option(ssc.checkpointDir)
  val serializableHadoopConf =
    new SerializableConfiguration(ssc.sparkContext.hadoopConfiguration)

  // Function to start the receiver on the worker node
  val startReceiverFunc: Iterator[Receiver[_]] => Unit =
    (iterator: Iterator[Receiver[_]]) => {
      if (!iterator.hasNext) {
        throw new SparkException(
          "Could not start receiver as object not found.")
      }
      if (TaskContext.get().attemptNumber() == 0) {
        val receiver = iterator.next()
        assert(iterator.hasNext == false)
        val supervisor = new ReceiverSupervisorImpl(
          receiver, SparkEnv.get, serializableHadoopConf.value, checkpointDirOption)
        supervisor.start()
        supervisor.awaitTermination()
      } else {
        // It's restarted by TaskScheduler, but we want to reschedule it again. So exit it.
      }
    }

  // Create the RDD using the scheduledLocations to run the receiver in a Spark job
  val receiverRDD: RDD[Receiver[_]] =
    if (scheduledLocations.isEmpty) {
      ssc.sc.makeRDD(Seq(receiver), 1)
    } else {
      val preferredLocations = scheduledLocations.map(_.toString).distinct
      ssc.sc.makeRDD(Seq(receiver -> preferredLocations))
    }
  receiverRDD.setName(s"Receiver $receiverId")
  ssc.sparkContext.setJobDescription(s"Streaming job running receiver $receiverId")
  ssc.sparkContext.setCallSite(Option(ssc.getStartSite()).getOrElse(Utils.getCallSite()))

  val future = ssc.sparkContext.submitJob[Receiver[_], Unit, Unit](
    receiverRDD, startReceiverFunc, Seq(0), (_, _) => Unit, ())
  // We will keep restarting the receiver job until ReceiverTracker is stopped
  future.onComplete {
    case Success(_) =>
      if (!shouldStartReceiver) {
        onReceiverJobFinish(receiverId)
      } else {
        logInfo(s"Restarting Receiver $receiverId")
        self.send(RestartReceiver(receiver))
      }
    case Failure(e) =>
      if (!shouldStartReceiver) {
        onReceiverJobFinish(receiverId)
      } else {
        logError("Receiver has been stopped. Try to restart it.", e)
        logInfo(s"Restarting Receiver $receiverId")
        self.send(RestartReceiver(receiver))
      }
  }(submitJobThreadPool)
  logInfo(s"Receiver ${receiver.streamId} started")
}

看完上述內容,你們對如何理解Receiver啟動以及啟動源碼分析有進一步的了解嗎?如果還想了解更多知識或者相關內容,請關注億速云行業資訊頻道,感謝大家的支持。

向AI問一下細節

免責聲明:本站發布的內容(圖片、視頻和文字)以原創、轉載和分享為主,文章觀點不代表本網站立場,如果涉及侵權請聯系站長郵箱:is@yisu.com進行舉報,并提供相關證據,一經查實,將立刻刪除涉嫌侵權內容。

AI

蓝山县| 广南县| 双流县| 四子王旗| 德钦县| 敖汉旗| 新乡市| 宝兴县| 高雄市| 广平县| 桃园县| 元朗区| 永靖县| 雷山县| 台前县| 海门市| 女性| 石阡县| 仁化县| 兰溪市| 化隆| 南溪县| 安阳县| 天津市| 都匀市| 嘉义县| 上高县| 西畴县| 迁西县| 山东省| 达州市| 铁岭市| 莱西市| 东明县| 浪卡子县| 龙井市| 凤庆县| 板桥市| 左云县| 台安县| 宝山区|