Hadoop 类AggregateWordCount源代码注释-humengez-ChinaUnix博客

codercodercoder！

首页　| 　博文目录　| 　关于我

humengez

博客访问： 83056
博文数量： 29
博客积分： 0
博客等级：民兵
技术积分： 225
用户组：普通用户
注册时间： 2014-03-06 15:31

文章分类

全部博文（29）

经典算法（2）
JAVA（11）
linux（6）

netlink（6）
hadoop（4）
network（6）
未分配的博文（0）

文章存档

2015年（18）

2014年（11）

我的朋友

相关博文

Hadoop 类AggregateWordCount源代码注释

分类： HADOOP

2014-09-03 11:04:52

转自http://blog.csdn.net/wangxw8746/article/details/9230323

点击(此处)折叠或打开

package org.apache.hadoop.examples;
import java.io.IOException;
import java.util.ArrayList;
import java.util.StringTokenizer;
import java.util.Map.Entry;
import org.apache.hadoop.io.Text;
import org.apache.hadoop.mapred.JobClient;
import org.apache.hadoop.mapred.JobConf;
import org.apache.hadoop.mapred.lib.aggregate.ValueAggregatorBaseDescriptor;
import org.apache.hadoop.mapred.lib.aggregate.ValueAggregatorJob;
/**
这个是hadoop的map/reduce的例子，是对例子WordCount利用系统已经实现的map/reduce类进行简化。系统已经实现的ValueAggregatorBaseDescriptor 和ValueAggregatorJob已经实现各种数据类型的求和最大值，最小值的算法。类型如下：
UniqValueCount
LongValueSum
DoubleValueSum
ValueHistogram
LongValueMax
LongValueMin
StringValueMax
StringValueMin
具体请看相关的源代码。
这个job的执行必须用-jarlibs执行，不然会报configured错误。
执行命令如下：
hadoop jar hadoop-example.jar -libjars hadoop-example.jar shakepoems.text out_aggregate_his 3 textinputformat
* This is an example Aggregated Hadoop Map/Reduce application. It reads the
* text input files, breaks each line into words and counts them. The output is
* a locally sorted list of words and the count of how often they occurred.
*
* To run: bin/hadoop jar hadoop-*-examples.jar aggregatewordcount in-dir
* out-dir numOfReducers textinputformat
*
*/
public class AggregateWordCount {
/*继承类ValueAggregatorBaseDescriptor */
public static class WordCountPlugInClass extends
ValueAggregatorBaseDescriptor {
@Override
public ArrayList<Entry<Text, Text>> generateKeyValPairs(Object key,
Object val) {
String countType = LONG_VALUE_SUM;//指定算法类型是long类型的求和
ArrayList<Entry<Text, Text>> retv = new ArrayList<Entry<Text, Text>>();
String line = val.toString();
StringTokenizer itr = new StringTokenizer(line);
while (itr.hasMoreTokens()) {
Entry<Text, Text> e = generateEntry(countType, itr.nextToken(), ONE);
if (e != null) {
retv.add(e);
}
}
return retv;
}
}
/**用静态类ValueAggregatorJob执行job
* The main driver for word count map/reduce program. Invoke this method to
* submit the map/reduce job.
*
* @throws IOException
* When there is communication problems with the job tracker.
*/
@SuppressWarnings("unchecked")
public static void main(String[] args) throws IOException {
JobConf conf = ValueAggregatorJob.createValueAggregatorJob(args
, new Class[] {WordCountPlugInClass.class});
JobClient.runJob(conf);
}
}

阅读(1984) | 评论(0) | 转发(0) |

上一篇：hadoop wordmean源码及注释

下一篇：Hadoop WordCount解读

给主人留下些什么吧！~~

感谢所有关心和支持过ChinaUnix的朋友们

16024965号-6