問(wèn)題內(nèi)容
我是 spark 新手。我試圖展平數(shù)據(jù)框,但未能通過(guò)“爆炸”做到這一點(diǎn)。
原始數(shù)據(jù)框架構(gòu)如下:
id|approvaljson 1|[{"approvertype":"1st line manager","status":"approved"},{"approvertype":"2nd line manager","status":"approved"}] 2|[{"approvertype":"1st line manager","status":"approved"},{"approvertype":"2nd line manager","status":"rejected"}]
登錄后復(fù)制
我需要將其轉(zhuǎn)換為以下架構(gòu)?
id|approvaltype|status 1|1st line manager|approved 1|2nd line manager|approved 2|1st line manager|approved 2|2nd line manager|rejected
登錄后復(fù)制
我已經(jīng)嘗試過(guò)
df_exploded = df.withcolumn("approvaljson", explode("approvaljson"))
登錄后復(fù)制
但是我得到了錯(cuò)誤:
Cannot resolve "explode(ApprovalJSON)" due to data type mismatch: parameter 1 requires ("ARRAY" or "MAP") type, however, "ApprovalJSON" is of "STRING" type.;
登錄后復(fù)制
正確答案
首先將類(lèi)似 json 的字符串解析為結(jié)構(gòu)數(shù)組,然后使用 inline
將數(shù)組分解為行和列
df1 = df.withcolumn("approvaljson", f.from_json("approvaljson", schema="array")) df1 = df1.select("id", f.inline('approvaljson'))
登錄后復(fù)制
結(jié)果
df1.show() +---+----------------+--------+ | ID| ApproverType| Status| +---+----------------+--------+ | 1|1st Line Manager|Approved| | 1|2nd Line Manager|Approved| | 2|1st Line Manager|Approved| | 2|2nd Line Manager|Rejected| +---+----------------+--------+
登錄后復(fù)制